How to Use Elasticsearch Aggregations (With Examples)
Use Elasticsearch aggregations for faceted search, metrics, and time-series analytics. Step-by-step examples for terms, date_histogram, range, and composite.
Overview
Any time you need counts, sums, or percentiles from an Elasticsearch index, aggregations are the obvious place to start. They run directly on the inverted index, which keeps counts and metrics over millions of documents fast enough for live facets and dashboards.
A single request can combine bucket aggregations that split documents into groups with metric aggregations that compute values inside each group. Nest them to build time-series, percentiles, and faceted summaries without shipping data to a separate batch job.
I’ve used this pattern on product catalogs where the same request returns matching products and facet counts for category, brand, and price range. Without aggregations, you would need a second query or a batch job, and the facets would be stale by the time they reach the UI.
Most examples target Elasticsearch 8.x. They also work on 7.10+, but
date_histogram switched from interval to calendar_interval between major
versions. The official docs keep a complete aggregation reference at
Elasticsearch Aggregations.
The request cache and eager global ordinals can speed up aggregations, but only after the query itself is well structured. I cover both in Best Practices.
When to Use
- You’re building faceted search with category-level count filters. The query side is covered in Full-Text Search.
- Your analytics dashboards need sub-second numbers across large document sets.
- You want to bucket time series and nest statistics like sum, average, or percentiles.
- You need unique counts, top hits per bucket, or derivative metrics in the same request.
When to avoid
- The query looks like a SQL join across several tables: aggregations stay inside one index and can’t span them.
- The field you want to aggregate isn’t indexed, or it’s a tokenized
textfield with no.keywordsubfield. - You need exact counts over very high-cardinality fields; prefer
compositeor tuneshard_sizeinstead of a plaintermsaggregation.
Solution
Terms aggregation for faceted search
GET /products/_search
{
"size": 0,
"aggs": {
"categories": {
"terms": {
"field": "category.keyword",
"size": 10
}
}
}
}
Faceted search with the JavaScript client
// client/SearchClient.ts
async function getCategoryFacets(query: string) {
const response = await client.search({
index: 'products',
size: 0,
query: { match: { name: query } },
aggs: {
categories: {
terms: { field: 'category.keyword', size: 20 }
},
brands: {
terms: { field: 'brand.keyword', size: 20 }
}
}
});
return {
categories: response.aggregations?.categories.buckets,
brands: response.aggregations?.brands.buckets,
};
}
Date histogram with nested metrics
GET /orders/_search
{
"size": 0,
"aggs": {
"sales_over_time": {
"date_histogram": {
"field": "created_at",
"calendar_interval": "month"
},
"aggs": {
"revenue": {
"sum": { "field": "total_amount" }
},
"avg_order_value": {
"avg": { "field": "total_amount" }
}
}
}
}
}
Range aggregation for pricing tiers
GET /products/_search
{
"size": 0,
"aggs": {
"price_ranges": {
"range": {
"field": "price",
"ranges": [
{ "to": 50, "key": "budget" },
{ "from": 50, "to": 200, "key": "mid-range" },
{ "from": 200, "key": "premium" }
]
}
}
}
}
Composite aggregation for deep pagination
GET /events/_search
{
"size": 0,
"aggs": {
"events_by_region": {
"composite": {
"size": 100,
"sources": [
{ "region": { "terms": { "field": "region.keyword" } } },
{ "day": { "date_histogram": { "field": "timestamp", "calendar_interval": "day" } } }
]
}
}
}
}
async function paginateAggregations(afterKey = null) {
const body = {
size: 0,
aggs: {
events_by_region: {
composite: {
size: 100,
sources: [
{ region: { terms: { field: 'region.keyword' } } },
{ day: { date_histogram: { field: 'timestamp', calendar_interval: 'day' } } }
],
...(afterKey && { after: afterKey })
}
}
}
};
const response = await client.search({ index: 'events', body });
const { buckets, after_key } = response.aggregations.events_by_region;
if (after_key) {
console.log(`Got ${buckets.length} buckets, fetching next page...`);
return [...buckets, ...await paginateAggregations(after_key)];
}
return buckets;
}
Python client: terms, stats, and percentiles
from elasticsearch import Elasticsearch
es = Elasticsearch("http://localhost:9200")
response = es.search(
index="products",
size=0,
query={"match": {"name": "laptop"}},
aggs={
"categories": {
"terms": {"field": "category.keyword", "size": 20}
},
"price_stats": {
"stats": {"field": "price"}
},
"price_percentiles": {
"percentiles": {"field": "price", "percents": [25, 50, 75, 95]}
}
}
)
print(response["aggregations"]["categories"]["buckets"])
print(response["aggregations"]["price_stats"])
Filter bucket aggregation
GET /products/_search
{
"size": 0,
"aggs": {
"in_stock": {
"filter": { "term": { "status": "in_stock" } },
"aggs": {
"avg_price": { "avg": { "field": "price" } }
}
},
"out_of_stock": {
"filter": { "term": { "status": "out_of_stock" } },
"aggs": {
"avg_price": { "avg": { "field": "price" } }
}
}
}
}
Cardinality for unique counts
GET /orders/_search
{
"size": 0,
"aggs": {
"unique_customers": {
"cardinality": {
"field": "customer_id",
"precision_threshold": 40000
}
}
}
}
Cardinality uses HyperLogLog++ for approximate distinct counts. The
precision_threshold trades accuracy for memory: higher values are more
accurate but consume more heap. With precision_threshold: 40000, Elasticsearch returns counts within 1% of the true value. That margin is enough for most dashboards.
Top hits per bucket
GET /products/_search
{
"size": 0,
"aggs": {
"by_category": {
"terms": { "field": "category.keyword", "size": 10 },
"aggs": {
"top_products": {
"top_hits": {
"size": 3,
"sort": [{ "popularity": "desc" }],
"_source": ["name", "price", "rating"]
}
}
}
}
}
}
Pipeline aggregations
GET /orders/_search
{
"size": 0,
"aggs": {
"monthly_sales": {
"date_histogram": {
"field": "created_at",
"calendar_interval": "month"
},
"aggs": {
"revenue": {
"sum": { "field": "total_amount" }
},
"revenue_derivative": {
"derivative": { "buckets_path": "revenue" }
},
"revenue_moving_avg": {
"moving_avg": {
"buckets_path": "revenue",
"window": 3,
"model": "holt"
}
}
}
}
}
}
Mapping for keyword aggregation
PUT /products
{
"mappings": {
"properties": {
"category": {
"type": "text",
"fields": {
"keyword": { "type": "keyword" }
}
},
"brand": {
"type": "text",
"fields": {
"keyword": { "type": "keyword" }
}
}
}
}
}
This mapping is what makes the .keyword subfield examples work. Elasticsearch analyzes the text field for search, but stores the keyword
sibling as a single token for grouping and sorting.
Explanation
Bucket aggregations split documents into groups. terms and range are bucket
aggregations; date_histogram splits by time. Metric aggregations like sum,
avg, stats, and percentiles run inside each bucket.
Nesting aggregations lets you answer multi-level questions: monthly revenue per
category, average price per price range, or percentile latency per region. Set
size: 0 and Elasticsearch skips the hits, returning only aggregation results.
Skipping the hits is faster when the documents themselves aren’t needed.
Aggregations run in two main phases: a shard-level phase and a reduce phase. During the shard-level phase, each shard handles its own documents and computes a partial result. In the reduce phase, the
coordinating node merges those partials into the final response. This design is
why size: 0 is so effective: you skip the fetch and merge of the actual hits
and only ship the compact aggregation results.
The terms aggregation is approximate because it uses a per-shard priority
queue. Elasticsearch returns doc_count_error_upper_bound so you can judge how
much the count may be off. If exact counts matter, increase shard_size (the
default is the same as size) or switch to composite, which walks the
doc-values in sorted order.
For pagination, composite is the safest aggregation to use. It returns an after_key that you pass to the next request. Unlike terms with from, the after_key is
stable while new documents are being indexed. For the document side of paging,
see Cursor-Based Pagination with PostgreSQL.
Metric aggregations are generally exact, but cardinality isn’t. It uses
HyperLogLog++ to estimate distinct counts with a configurable
precision_threshold. At 40000, Elasticsearch keeps the error below 1%, which is usually fine for dashboard metrics. If you need exact unique counts, use a
terms aggregation with a large enough size, though it will consume more
memory.
Pipeline aggregations such as derivative and moving_avg are powerful but
costly. They run a second pass over the bucket list, so they push up CPU and memory
usage on large time ranges. I avoid them for dashboards with hundreds of
buckets and prefer to compute trends in the application layer when possible.
Text fields are analyzed and tokenized, so they can’t be aggregated directly.
For counts, groupings, and filters, use the .keyword subfield: that sibling is
stored as a single unanalyzed token, which is exactly what aggregations need.
post_filter applies search filters after aggregations are computed. Use it
when users filter results but you want to keep the original facet counts.
Variants
|| Aggregation | Use case | Key parameters |
|| --- | --- | --- |
|| terms | Count by category, brand, or status | field, size, shard_size |
|| date_histogram | Time-series bucketing | field, calendar_interval |
|| range | Predefined bands such as price tiers | ranges |
|| composite | Paginating over high-cardinality keys | sources, size, after |
|| cardinality | Approximate unique counts | precision_threshold |
|| top_hits | Best document per bucket | size, sort, _source |
|| filter | Conditional sub-aggregations | filter query |
Use this table as a quick reference, not as a replacement for the examples.
terms and date_histogram cover most day-to-day use cases, while composite
is the right choice when you need to stream buckets instead of returning a top-N.
Best Practices
- Set
size: 0when you only need aggregations and not the search hits. - Aggregate on
keywordsubfields, not on analyzedtextfields. - Switch to
compositewhen an aggregation could return more than a few thousand buckets. - Enable
eager_global_ordinalson fields you aggregate frequently, especially for high-cardinalityterms. - Use
post_filterto filter returned results without changing aggregation counts. - Tune
precision_thresholdoncardinalityto balance memory and accuracy. - Set a realistic
shard_sizefortermswhen counts need to be closer to exact; I usually start with5 * sizeand measure. - Use
min_doc_count: 1to drop empty buckets unless you actually need them. - Cache expensive aggregation requests with the request cache when the underlying data doesn’t change often.
Common Mistakes
- Aggregating on a
textfield instead of its.keywordsubfield. - Setting
size: 10000on atermsaggregation and assuming the cluster has unlimited heap. - Paginating large
termsresults withoutcomposite. - Forgetting that
termsandcardinalityreturn approximate counts. - Running heavy pipeline aggregations on very large time ranges.
- Using
fromon atermsaggregation and getting unstable results. - Ignoring
doc_count_error_upper_boundand treatingtermscounts as exact. - Running
top_hitswithout asort, which makes the returned document unpredictable.
See Also
- Full-Text Search — for the query side of the same use case.
- Complete Guide to Elasticsearch Cluster Setup — when you want to scale from a single node to production.
- Database Views and Materialized Views — an alternative to real-time aggregations when precomputed results are enough.
- Elasticsearch Aggregations — official reference.
- Composite aggregation — official docs for deep pagination.
- Cardinality aggregation —
details on
precision_threshold.
Frequently Asked Questions
Can I combine multiple aggregations in one query?
Yes. You can place several top-level aggregations and nest bucket and metric aggregations inside each other in the same request.
How do I filter results without changing aggregation counts?
Use post_filter when you want filters applied after the aggregations are
computed.
The aggregations see the full query, while the returned hits are filtered.
Are Elasticsearch aggregations exact on large datasets?
terms and cardinality are approximate, so don't expect exact counts.
Increase shard_size, use composite for exact counts, or raise
precision_threshold to improve cardinality accuracy.
Why should I use composite instead of terms for pagination?
composite returns a stable after_key and walks the result set in sorted
order. terms pagination with from isn't reliable because bucket order can
change as data is indexed.
What is the difference between filter and post_filter?
A filter aggregation adds a bucket under the aggregation tree. post_filter
applies only to the search hits, so the aggregation values stay the same.
How do I debug slow aggregation queries?
Start with the
Profile API.
It breaks each aggregation phase into collect, build_aggregation, and
reduce times, so you can see whether the slowness is at the shard or reduce
level. You can also enable slow_log for queries that exceed a threshold value.
Related Resources
Full-Text Search
How to implement full-text search with Elasticsearch, Meilisearch, and PostgreSQL.
RecipeCRUD Operations with MongoDB and Mongoose
How to perform Create, Read, Update, and Delete operations in MongoDB using Mongoose ODM with Node.js and Express
RecipeCreate and Use Database Views and Materialized Views
How to create and use database views and materialized views to simplify queries and improve read performance.
RecipeCursor-Based Pagination in PostgreSQL (Keyset vs OFFSET)
Implement efficient cursor-based pagination for large datasets in PostgreSQL, avoiding OFFSET performance degradation with indexed keyset pagination and stable sort ordering
GuideComplete Guide to Elasticsearch Cluster Setup
Deploy and scale Elasticsearch clusters. Covers node roles, sharding, replicas, index templates, mapping, snapshots, and production tuning for search at scale.
GuideFull-Text Search — Implement Search That Actually Works
A practical guide to full-text search: PostgreSQL tsvector, Elasticsearch indexing, query design, relevance tuning, and building search that users trust with autocomplete, faceting, and typo tolerance.