Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
elasticsearch

Understanding and Customizing Analyzers in Elasticsearch

Posted on October 6, 2025October 5, 2025 by admin

When you search for text in Elasticsearch, the system doesn’t just compare exact words.
It first processes your text, breaking it into tokens, lowercasing, removing stopwords, and sometimes even finding the root form of words.

This process is handled by something called an analyzer. Understanding analyzers is the first step to building powerful and accurate search features.

What Is an Analyzer?

An analyzer in Elasticsearch is like a mini pipeline that processes text before it’s stored or searched.

It usually has three parts:

PartPurposeExample
Character filterClean up text before tokenizingRemove HTML tags, replace symbols
TokenizerSplit text into tokens (words)Split by spaces or punctuation
Token filterModify tokensLowercase, remove stopwords, stem words

Example:

Input text:

"The Quick Brown Fox Jumps Over The Lazy Dog!"

Analyzer output:

["the", "quick", "brown", "fox", "jumps", "over", "the", "lazy", "dog"]

Why Analyzers Matter

Choosing the right analyzer affects:

  • How text is indexed (stored)
  • How search queries are processed
  • Whether users get relevant results

For example:

  • A keyword analyzer will only match exact text.
  • A language analyzer can understand word forms like “run”, “running”, “ran”.
  • A custom analyzer can remove symbols or accents.

So, if you choose the wrong analyzer, your search results can become incomplete or noisy.

Common Built-In Analyzers

Elasticsearch provides several ready-to-use analyzers for common use cases.

Elasticsearch provides several ready-to-use analyzers for common use cases.

AnalyzerDescriptionExample Output
standardDefault analyzer, splits by word boundaries"The Quick Fox" → ["the", "quick", "fox"]
simpleSplits by non-letters, lowercase"[email protected]" → ["john", "doe", "example", "com"]
whitespaceSplits only on spaces"AI_MachineLearning NLP" → ["AI_MachineLearning", "NLP"]
stopRemoves common stopwords like “a”, “the”, “is”"The fox is red" → ["fox", "red"]
keywordKeeps entire text as one token"SKU-12345" → ["SKU-12345"]
language (e.g., english, indonesian)Stemming + stopwords by language"berlari" → "lari"

Testing an Analyzer

You can test any analyzer with the _analyze API.

POST _analyze
{
  "analyzer": "standard",
  "text": "The Quick Brown Fox!"
}

Output:

{
  "tokens": [
    { "token": "the" },
    { "token": "quick" },
    { "token": "brown" },
    { "token": "fox" }
  ]
}

Tip: Always test your analyzer before applying it to a big dataset.
It’s the easiest way to see how Elasticsearch “understands” your text.

See also  Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana

Language-Specific Analyzers

Elasticsearch supports many languages — each with its own stopwords and stemming rules.

Example: English

POST _analyze
{
  "analyzer": "english",
  "text": "running runners run"
}

Output:

["run", "runner", "run"]

Example: Indonesian

POST _analyze
{
  "analyzer": "indonesian",
  "text": "Mahasiswa sedang belajar di universitas"
}

Output:

["mahasiswa", "belajar", "universitas"]

The analyzer automatically removes common words and keeps root forms.

Related posts:

Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search

Basic Concept of ElasticSearch (Part 1): Introduction

Maintaining Super Large Datasets in Elasticsearch

Pages: 1 2
Category: ElasticSearch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme