Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
elasticsearch

Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type

Posted on October 6, 2025October 5, 2025 by admin

Reverse Analyzer, Suffix or “Ends With” Search

By default, Elasticsearch matches from the start of words.
But what if you need to match from the end, for example, file extensions like .pdf or .zip?

Use a reverse token filter to flip the text, and then apply Edge N-Gram.

Example Mapping

PUT reverse_index
{
  "settings": {
    "analysis": {
      "filter": {
        "edge_reverse": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 10
        }
      },
      "analyzer": {
        "reverse_autocomplete": {
          "tokenizer": "standard",
          "filter": ["lowercase", "reverse", "edge_reverse", "reverse"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "filename": {
        "type": "text",
        "analyzer": "reverse_autocomplete",
        "search_analyzer": "standard"
      }
    }
  }
}

Now, searching for “pdf” or “zip” will match “report.pdf” or “backup.zip”. Use for suffix-based search like extensions, domain endings, or last words.

Fuzzy Search, Handling Typos

Users often make spelling mistakes.
Fuzzy search helps Elasticsearch find results even with small typos or missing letters.

Concept

Fuzzy search uses Levenshtein edit distance, how many edits are needed to turn one word into another.

Examples:

  • “iphon” → 1 edit away from “iphone”
  • “googel” → 2 edits away from “google”

Example Query

GET products/_search
{
  "query": {
    "match": {
      "name": {
        "query": "iphon",
        "fuzziness": "AUTO"
      }
    }
  }
}

Matches “iPhone”, even though user typed “iphon”. You can also set fuzziness manually:

"fuzziness": 2

Performance Tip

Use fuzzy search only for short fields (titles, names),
because fuzzy queries are slower on long text fields.

Search-as-you-type — Built-In Autocomplete Field

From Elasticsearch 7.x onward, there’s an easier way:
the search_as_you_type field type — no need to define analyzers manually.

This field automatically creates small subfields (_2gram, _3gram, _index_prefix)
so it behaves like edge n-gram but optimized.

Example Mapping

PUT product_titles
{
  "mappings": {
    "properties": {
      "title": { "type": "search_as_you_type" }
    }
  }
}

Example Query

GET product_titles/_search
{
  "query": {
    "multi_match": {
      "query": "iph",
      "type": "bool_prefix",
      "fields": [
        "title",
        "title._2gram",
        "title._3gram",
        "title._index_prefix"
      ]
    }
  }
}

Matches “iPhone 15 Pro Max” as you type “iph”. Works with fuzziness: 1 for typo tolerance.

See also  Maintaining Super Large Datasets in Elasticsearch

Combining Them for Best Experience

FeatureGoalIdeal Analyzer / Query
Autocomplete (prefix)Find as user typesEdge N-Gram / search_as_you_type
Suffix searchMatch ending (e.g. .pdf)Reverse Edge N-Gram
Typo toleranceHandle spelling errorsFuzzy search (fuzziness: AUTO)
Mid-word matchSearch within wordsN-Gram
Phrase matchKeep ordermatch_phrase or match_phrase_prefix

Example: All-In-One Setup

"mappings": {
  "properties": {
    "title": {
      "type": "text",
      "fields": {
        "autocomplete": { "type": "text", "analyzer": "autocomplete" },
        "reverse": { "type": "text", "analyzer": "reverse_autocomplete" },
        "keyword": { "type": "keyword" }
      }
    }
  }
}

Then use queries like:

{
  "bool": {
    "should": [
      { "match": { "title.autocomplete": "iphon" } },
      { "match": { "title.reverse": "pdf" } },
      { "match": { "title": { "query": "iphon", "fuzziness": 1 } } }
    ]
  }
}

Covers prefix, suffix, and typo in one unified search pipeline.

Related posts:

Elasticsearch Best Practices for Beginners

Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades

Basic Concept of ElasticSearch (Part 3): Translog, Flush, and Refresh

Pages: 1 2 3
Category: ElasticSearch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme