Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
elasticsearch

Understanding and Customizing Analyzers in Elasticsearch

Posted on October 6, 2025October 5, 2025 by admin

Creating a Custom Analyzer

Sometimes, you need more control — maybe you want to:

  • Remove accents (café → cafe)
  • Add synonyms
  • Use lowercase only
  • Ignore specific words

Then you can define a custom analyzer inside index settings.

Example:

PUT my_index
{
  "settings": {
    "analysis": {
      "analyzer": {
        "my_custom_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": ["lowercase", "stop", "asciifolding"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": { "type": "text", "analyzer": "my_custom_analyzer" }
    }
  }
}

How It Works:

  1. Tokenizer: breaks text into words.
  2. Lowercase filter: converts all to lowercase.
  3. Stop filter: removes common words (“the”, “a”).
  4. Asciifolding: removes accents (café → cafe).

Now when you index:

POST my_index/_doc
{ "title": "Café in Jakarta" }

and search for "cafe", it will match perfectly — even without the accent.

Adding Synonyms

You can also teach Elasticsearch that some words mean the same thing.
For example: “car”, “auto”, “vehicle”.

PUT synonym_index
{
  "settings": {
    "analysis": {
      "filter": {
        "synonym_filter": {
          "type": "synonym",
          "synonyms": [
            "car, automobile, vehicle"
          ]
        }
      },
      "analyzer": {
        "my_synonym_analyzer": {
          "tokenizer": "standard",
          "filter": ["lowercase", "synonym_filter"]
        }
      }
    }
  }
}

Now searching "automobile" will also find "car" and "vehicle".

Synonyms make search more human, users don’t always type the same word you used.

Comparison: Standard vs Custom Analyzer

FeatureStandard AnalyzerCustom Analyzer
PredefinedYesNo
Easy to useYesNeeds config
Language supportBasicFlexible
SynonymsNoYes
FiltersLimitedAny combination
Best forGeneral textDomain-specific search

Analyzer Best Practices

  • Always lowercase your text: case differences don’t matter in most searches.
  • Use language analyzers when your data has meaningful word forms.
  • Add stopwords carefully: removing too many words may break phrases.
  • Test with _analyze before indexing large data.
  • Use search_analyzer if you want a different analyzer at query time.
See also  Basic Concept of ElasticSearch (Part 3): Translog, Flush, and Refresh

Example:

"analyzer": "my_custom_analyzer",
"search_analyzer": "standard"

Indexes text using custom rules but searches more broadly.

How Analyzers Affect Search and Autocomplete

Even though analyzers seem low-level, they’re the foundation of autocomplete and fuzzy matching.

For example:

  • An analyzer with lowercase + edge n-gram: enables “search as you type”
  • A language analyzer: improves semantic search
  • A custom analyzer with synonyms: improves query recall

So, understanding analyzers = understanding how Elasticsearch finds what users mean, not just what they type.

Conclusion

Analyzers are the brain of Elasticsearch text search. They define how your text is split, normalized, and matched, and with a bit of customization, they can make your search engine smarter than ever.

“If mapping is the skeleton, analyzer is the brain.”

Start with the built-in ones (standard, english, indonesian),
and once you understand how they tokenize text,
experiment with your own analyzer to match your domain’s language and behavior.

Related posts:

Maintaining Super Large Datasets in Elasticsearch

Basic Concept of ElasticSearch (Part 1): Introduction

Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search

Pages: 1 2
Category: ElasticSearch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme