Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
elasticsearch

Basic Concept of ElasticSearch (Part 1): Introduction

Posted on June 29, 2021 by admin

Elasticsearch is a text search engine based on Lucene. Basically Lucene is a plugin written 100% in Java that have powerfull capability to search text with extremely fast. ElasticSearch wraps it so that we can utilize it over HTTP call with JSON interface. Besides that, it has ability to distribute documents across multiple shards in a cluster.

Instead of search text directly, ElasticSearch converts it into an index that contains terms frequency and also its location. In this article, I will show you how ElasticSearch works and why ElasticSearch has good performance for text searching. I divide this into some parts so that it is easier to read.

Comparison Between ElasticSearch and SQL

Even though ElasticSearch and SQL are not apple to apple comparison. I can give you an analogy between terminology in ElasticSearch and SQL.

ElasticSearchSQL
IndexDatabase
DocumentsRows
PropertiesColumns

Before ElasticSearch 7, there is another terminology in ElasticSearch called as Type which is grouping documents that have the same structure and you can imagine it as Table in SQL. But after ElasticSearch 7, Type has been removed. Because, fields that have the same name in different doc types are stored in the same Lucene field internally. It makes those fields are dependent implicitly. While in SQL, tables are independent each other. Thus, there is no reason to have Type, and use different Index if needed.

Terminologies in ElasticSearch

Index

Index is a container to store data similar to a database in the relational databases. An index contains a collection of documents that have similar characteristics or are logically related.

See also  Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades

If we take an example of an e-commerce website, there will be one index for products, one for customers, and so on. Indices are identified by the lowercase name. The index name is required to perform the add, update, or delete operations on the document.

Documents

Document is the piece indexed by Elasticsearch and represented in the JSON format. We can add as many documents as we want into an index. The following snippet shows how to create a document of type mobile in the index store. 

Property or Field

List of fields in the document. Fields also carry the data type information with them. Data Types are similar to what we see in any other programming language. We also able to apply custom analyzer for each field. For example of an e-commerce website and the index is product, then field can be name, tag, category, image, price, etc.

Meta Fields

Meta fields store additional information about the document. Meta fields are meant for mostly internal usage purpose. Meta field names start with an underscore, example: _shards, _index, _type, _id, _score, _routing, etc

Relation Between ElasticSearch Index and Lucene Index

Relation Between ElasticSearch Index and Lucene Index

As I mention before, ElasticSearch is built over Lucene as its core. The biggest component of ElasticSearch is index. For each index it can have one or more shards that are distributed across some nodes.

ElasticSearch shard is Lucene index. Each Lucene index is split into some chunks called segments (or we can call it as Lucene segments). Each segment has its own mini index called inverted index and also stores original data by default. Segments are loaded in memory but it can persist in disk during flush time.

See also  Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type

ElasticSearch utilize 2 kind of memory allocations, on heap (JVM Heap) and off heap (filesystem cache / page cache and other usage). JVM Heap is controlled by ElasticSearch, while page cache is controlled by Operating System. Page cache is used for quicker access for certain content. At initialization, we have to define percentage of memory allocations in JVM option’s file.

It is not recommended to allocate high JVM heap and the opposite. Usually we allocate 50% for JVM heap of total memory. However, we have to avoid allocate JVM memory over 32GB. HotSpot JVM uses a trick to compress object pointers when heaps are less than around 32GB. If we allocate more than 32GB, object pointers occupy double the space, and less memory will be available for operations, which eventually results in performance degradation.

ElasticSearch Data Type

Nowdays ElasticSearch supports many data types and grouped based on these types:

  • common types: binary, boolean, keywords, numbers, dates, alias
  • object and relational types: object, flattened, nested, join
  • structured data types: range, ip, version, murmur3
  • aggregate data types: aggregate_metric_double, histogram
  • text search types: text, annotated-text, completion, search_as_ypu_type, token_count
  • document ranking types: dense_vector, sparse_vector, rank_feature, rank_features
  • spatial data types: geo_point, geo_shape, point, shape
  • other types: percolator

You can find its detail on this page.

Conclusion

Elasticsearch is a text search engine based on Lucene. So that ElasticSearch has powerfull capability to search text over HTTP call with extremely fast.

Instead of search text directly, ElasticSearch converts it into an index that contains terms frequency and also its location. That is why it is called with inverted index, raw document is converted to an index, then the index refers back to the original document.

See also  Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple

In the next article, I will give you a brief of ElasticSearch architectural perspective.

Related posts:

Basic Concept of ElasticSearch (Part 2): Architectural Perspective

Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type

Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple

Category: ElasticSearch

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme