Integrating Apache Lucene for Bean Search — Part 5: Facets

Tuesday, Sep 15, 2026 | 3 minute read

David Pilato
Integrating Apache Lucene for Bean Search — Part 5: Facets

This post is part of a series:

Part 3 can already filter (genre:Club). A filter panel still needs something else: how many tracks sit in Club vs Techno under the current query.

You declared lucene-facet in Part 1. Counts live on the same in-process index as search. Lucene remains a derived cache: rebuild or upsert after writes, then recount.

Index the category

genre.raw is a StringField so genre:Club can be a term query. Counting is a different access pattern: you want a column of labels, not a stored field on each hit.

Add a SortedSetDocValuesFacetField next to the keyword (same idea for artist, album, key). Do not facet on a TextField — tokens are not checkbox labels. Numerics (bpm, rating, year) already carry DocValues from Part 1; range and value counts read those, no extra field type.

String genre = name(t.genre());
doc.add(new StringField(TrackIndexFields.GENRE_RAW, genre, Field.Store.YES));
// Category label for lucene-facet — not a replacement for the StringField above
doc.add(new SortedSetDocValuesFacetField("genre", genre));

SortedSetDocValuesFacetField is a helper. It does not write DocValues by itself. Wrap every add / upsert with FacetsConfig.build or you silently count nothing:

private static final FacetsConfig FACETS = new FacetsConfig();

private static Document indexedDocument(Track track) throws IOException {
    return FACETS.build(TrackDocumentMapper.toDocument(track));
}

Use that same FacetsConfig instance at search time. After commit, open a new NRT reader before counting.

Count

Lucene 10.5 fills a collector, then SortedSetDocValuesFacetCounts turns it into histograms. You do not walk ScoreDocs.

SortedSetDocValuesReaderState state =
        new DefaultSortedSetDocValuesReaderState(reader, FACETS);
FacetsCollector fc = FacetsCollectorManager.search(
                searcher, new MatchAllDocsQuery(), 1, new FacetsCollectorManager())
        .facetsCollector();
Facets facets = new SortedSetDocValuesFacetCounts(state, fc);
// facets.getAllChildren("genre") → Club=2, Techno=1

n=1 is enough: we want the collector, not a hit list (SearchService already resolved beans).

Drill-down and sideways

Under genre:Club, the key histogram should shrink. The genre panel should still show Techno — otherwise the user cannot change genre without clearing q.

Split the bookmarkable string: remainder (free text) is the base query; panel tokens become DrillDownQuery.add(dimension, query). Then DrillSideways runs one collector for the filtered set and one per selected dimension without that dimension’s own constraint:

Query base = TrackLuceneQueryBuilder.build(remainder); // no genre: / bpm: tokens
DrillDownQuery drillDown = new DrillDownQuery(FACETS, base);
drillDown.add("genre", TrackLuceneQueryBuilder.build("genre:Club"));

Facets luceneFacets = new DrillSideways(searcher, FACETS, state)
        .search(drillDown, 1)
        .facets;
PanelVisible buckets (count > 0)
genreClub and Techno
bpm80-90 and 120-130

The table is still Club ∩ 120–129. Only the counts omit that panel’s own filter — the usual e-commerce “narrow by brand without hiding the other brands”.

Keyword dims use SortedSetDocValuesFacetCounts. BPM bins use DoubleRangeFacetCounts on the existing bpm field (half-open ranges, always render every bucket including zeros). Rating / year use LongValueFacetCounts. If you need both on the same collectors, override DrillSideways.buildFacetsResult and wrap each collector with a MultiFacets — the default sideways class assumes one implementation.

Hand (value, count) to the UI

public record FacetBucket(String value, int count) {}

Map LabelAndValue once in TrackFacetService. The template prints Club (12). Clicking a checkbox still toggles a token in q; the next request filters the table (Part 3) and refreshes every panel here.

What this model does not do

One JVM, one Directory, rebuilt at startup. No replica, no cluster. Restart without a rebuild and search and facets are empty. If an upsert fails, fall back to a full rebuild so the cache cannot drift.

When the corpus or the ops model outgrows a process-local Lucene cache, the next step is a search server in front of the same beans — same q, same filter panel, a different engine behind TrackSearchIndex. That is a switch, not a rewrite of Parts 1–5.

© 2010 - 2026 David Pilato

Search is powered by Pagefind. Just hit CTRL+K or CMD+K to start searching.

Powered by Hugo with Dream and Devrel themes.

Details

I discovered Elasticsearch project in 2011. After contributed to the project and created open source plugins for it, David joined elastic the company in 2013 where he is Developer and Evangelist. He also created and still actively managing the French spoken language User Group. At elastic, he mainly worked on Elasticsearch source code, specifically on open-source plugins. In his free time, he likes talking about elasticsearch in conferences or in companies (Brown Bag Lunches AKA BBLs). He is also author of FSCrawler project which helps to index your pdf, open office, whatever documents in elasticsearch using Apache Tika behind the scene.

Who am I?

Developer | Evangelist at elastic and creator of the Elastic French User Group. Frequent speaker about all things Elastic, in conferences, for User Groups and in companies with BBL talks. In my free time, I enjoy coding and deejaying as DJ Elky, just for fun. Living with my children in Cergy, France.

Social Links