Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
elasticsearch

Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type

Posted on October 6, 2025October 5, 2025 by admin

Modern search systems don’t just find exact matches, they understand partial words, typos, and phrases as you type. Elasticsearch makes this possible with a mix of analyzers and special queries like N-Gram, Reverse, Fuzzy, and Search-as-you-type.

In this article, we’ll learn how these techniques work, when to use each, and how to combine them for a fast and smart search experience.

Why “Advanced” Text Search?

In real-world apps, users rarely type exact words.
They type fast, make mistakes, or expect results before finishing the query.

For example:

  • “iph”, should match “iPhone 15 Pro”
  • “iphon”, should still match “iPhone” (even with typo)
  • “.pdf”, should find “report.pdf”

To make that work, Elasticsearch uses special analyzers and queries under the hood.

N-Gram Analyzer — Matching Inside Words

The N-Gram analyzer splits text into small overlapping pieces called n-grams.
This allows search to match any part of a word, not just the beginning.

Example

Text: "search"
With min_gram: 3, max_gram: 5, tokens become:

["sea", "ear", "arc", "rch"]

So a user typing “arc” still finds “search”.

Example Mapping

PUT ngram_index
{
  "settings": {
    "analysis": {
      "tokenizer": {
        "my_ngram": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 5
        }
      },
      "analyzer": {
        "my_ngram_analyzer": {
          "tokenizer": "my_ngram",
          "filter": ["lowercase"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "name": { "type": "text", "analyzer": "my_ngram_analyzer" }
    }
  }
}

Now when you search:

GET ngram_index/_search
{
  "query": { "match": { "name": "arc" } }
}

It matches “search”, “arctic”, “arcade”, etc.

Trade-off

N-Gram creates many tokens, bigger index size.
Use it only for short fields like names or titles, not large documents.

Edge N-Gram, Perfect for Autocomplete

Edge N-Gram works like N-Gram, but it only creates tokens from the start of the word.
This makes it ideal for prefix matching (autocomplete).

See also  Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana

Text: "search"
Edge N-Gram tokens (min_gram=2, max_gram=5):

["se", "sea", "sear", "searc"]

So when a user types “sear”, it matches “search”.

Example Mapping

PUT edge_index
{
  "settings": {
    "analysis": {
      "tokenizer": {
        "edge_tokenizer": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 10
        }
      },
      "analyzer": {
        "autocomplete": {
          "tokenizer": "edge_tokenizer",
          "filter": ["lowercase"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "autocomplete",
        "search_analyzer": "standard"
      }
    }
  }
}

Now, as you type:

  • “se”: matches “search”, “service”, “secure”
  • “serv”: matches “server”, “service”

Great for instant autocomplete search bars.

Related posts:

Elasticsearch Best Practices for Beginners

Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades

Basic Concept of ElasticSearch (Part 3): Translog, Flush, and Refresh

Pages: 1 2 3
Category: ElasticSearch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme