Creating a Custom Analyzer
Sometimes, you need more control — maybe you want to:
- Remove accents (café → cafe)
- Add synonyms
- Use lowercase only
- Ignore specific words
Then you can define a custom analyzer inside index settings.
Example:
PUT my_index
{
"settings": {
"analysis": {
"analyzer": {
"my_custom_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "stop", "asciifolding"]
}
}
}
},
"mappings": {
"properties": {
"title": { "type": "text", "analyzer": "my_custom_analyzer" }
}
}
}
How It Works:
- Tokenizer: breaks text into words.
- Lowercase filter: converts all to lowercase.
- Stop filter: removes common words (“the”, “a”).
- Asciifolding: removes accents (café → cafe).
Now when you index:
POST my_index/_doc
{ "title": "Café in Jakarta" }
and search for "cafe", it will match perfectly — even without the accent.
Adding Synonyms
You can also teach Elasticsearch that some words mean the same thing.
For example: “car”, “auto”, “vehicle”.
PUT synonym_index
{
"settings": {
"analysis": {
"filter": {
"synonym_filter": {
"type": "synonym",
"synonyms": [
"car, automobile, vehicle"
]
}
},
"analyzer": {
"my_synonym_analyzer": {
"tokenizer": "standard",
"filter": ["lowercase", "synonym_filter"]
}
}
}
}
}
Now searching "automobile" will also find "car" and "vehicle".
Synonyms make search more human, users don’t always type the same word you used.
Comparison: Standard vs Custom Analyzer
| Feature | Standard Analyzer | Custom Analyzer |
|---|---|---|
| Predefined | Yes | No |
| Easy to use | Yes | Needs config |
| Language support | Basic | Flexible |
| Synonyms | No | Yes |
| Filters | Limited | Any combination |
| Best for | General text | Domain-specific search |
Analyzer Best Practices
- Always lowercase your text: case differences don’t matter in most searches.
- Use language analyzers when your data has meaningful word forms.
- Add stopwords carefully: removing too many words may break phrases.
- Test with
_analyzebefore indexing large data. - Use
search_analyzerif you want a different analyzer at query time.
Example:
"analyzer": "my_custom_analyzer", "search_analyzer": "standard"
Indexes text using custom rules but searches more broadly.
How Analyzers Affect Search and Autocomplete
Even though analyzers seem low-level, they’re the foundation of autocomplete and fuzzy matching.
For example:
- An analyzer with lowercase + edge n-gram: enables “search as you type”
- A language analyzer: improves semantic search
- A custom analyzer with synonyms: improves query recall
So, understanding analyzers = understanding how Elasticsearch finds what users mean, not just what they type.
Conclusion
Analyzers are the brain of Elasticsearch text search. They define how your text is split, normalized, and matched, and with a bit of customization, they can make your search engine smarter than ever.
“If mapping is the skeleton, analyzer is the brain.”
Start with the built-in ones (standard, english, indonesian),
and once you understand how they tokenize text,
experiment with your own analyzer to match your domain’s language and behavior.
