When you search for text in Elasticsearch, the system doesn’t just compare exact words.
It first processes your text, breaking it into tokens, lowercasing, removing stopwords, and sometimes even finding the root form of words.
This process is handled by something called an analyzer. Understanding analyzers is the first step to building powerful and accurate search features.
What Is an Analyzer?
An analyzer in Elasticsearch is like a mini pipeline that processes text before it’s stored or searched.
It usually has three parts:
| Part | Purpose | Example |
|---|---|---|
| Character filter | Clean up text before tokenizing | Remove HTML tags, replace symbols |
| Tokenizer | Split text into tokens (words) | Split by spaces or punctuation |
| Token filter | Modify tokens | Lowercase, remove stopwords, stem words |
Example:
Input text:
"The Quick Brown Fox Jumps Over The Lazy Dog!"
Analyzer output:
["the", "quick", "brown", "fox", "jumps", "over", "the", "lazy", "dog"]
Why Analyzers Matter
Choosing the right analyzer affects:
- How text is indexed (stored)
- How search queries are processed
- Whether users get relevant results
For example:
- A keyword analyzer will only match exact text.
- A language analyzer can understand word forms like “run”, “running”, “ran”.
- A custom analyzer can remove symbols or accents.
So, if you choose the wrong analyzer, your search results can become incomplete or noisy.
Common Built-In Analyzers
Elasticsearch provides several ready-to-use analyzers for common use cases.
Elasticsearch provides several ready-to-use analyzers for common use cases.
| Analyzer | Description | Example Output |
|---|---|---|
standard | Default analyzer, splits by word boundaries | "The Quick Fox" → ["the", "quick", "fox"] |
simple | Splits by non-letters, lowercase | "[email protected]" → ["john", "doe", "example", "com"] |
whitespace | Splits only on spaces | "AI_MachineLearning NLP" → ["AI_MachineLearning", "NLP"] |
stop | Removes common stopwords like “a”, “the”, “is” | "The fox is red" → ["fox", "red"] |
keyword | Keeps entire text as one token | "SKU-12345" → ["SKU-12345"] |
language (e.g., english, indonesian) | Stemming + stopwords by language | "berlari" → "lari" |
Testing an Analyzer
You can test any analyzer with the _analyze API.
POST _analyze
{
"analyzer": "standard",
"text": "The Quick Brown Fox!"
}
Output:
{
"tokens": [
{ "token": "the" },
{ "token": "quick" },
{ "token": "brown" },
{ "token": "fox" }
]
}
Tip: Always test your analyzer before applying it to a big dataset.
It’s the easiest way to see how Elasticsearch “understands” your text.
Language-Specific Analyzers
Elasticsearch supports many languages — each with its own stopwords and stemming rules.
Example: English
POST _analyze
{
"analyzer": "english",
"text": "running runners run"
}
Output:
["run", "runner", "run"]
Example: Indonesian
POST _analyze
{
"analyzer": "indonesian",
"text": "Mahasiswa sedang belajar di universitas"
}
Output:
["mahasiswa", "belajar", "universitas"]
The analyzer automatically removes common words and keeps root forms.
