Search your beans with Lucene — Index

Thursday, Sep 10, 2026 | 3 minute read

David Pilato
Search your beans with Lucene — Index

In the previous post we added Lucene to Maven, chose an analyzer, and mapped a Track bean to a search-ready Lucene Document. That is only half the story: you still need a small class that owns the index — and you should see what “inverted” means for a token like bob.

Own the index lifecycle

Wrap Lucene’s low-level types in a class dedicated to your bean. Create the directory and writer, rebuild or mutate, and close everything when the process shuts down. The playground does the same with an in-memory ByteBuffersDirectory — plain addDocument(doc), no facet rewrite yet:

// We will use an in-memory index
Directory dir = new ByteBuffersDirectory();

// Create the index writer with the analyzer
IndexWriter writer = new IndexWriter(dir, new IndexWriterConfig(analyzer));

// Create Lucene doc for track #255465792: Ultra Naté - Free (Bob Sinclar Remix)
Document doc255465792 = mapper.toDocument(Tracks.trackFrom(255465792));
writer.addDocument(doc255465792);

// Index track #172523747: Daft Punk - Around The World
Document doc172523747 = mapper.toDocument(Tracks.trackFrom(172523747));
writer.addDocument(doc172523747);

// Index track #106352474: Claude François - Cette année-là
Document doc106352474 = mapper.toDocument(Tracks.trackFrom(106352474));
writer.addDocument(doc106352474);

// Commit all the documents that have been indexed so far
writer.commit();

For a full library, wipe and reload under one write lock:

public void rebuild(List<Track> tracks) throws IOException {
    synchronized (writeLock) {
        writer.deleteAll();
        for (Track track : tracks) {
            writer.addDocument(TrackDocumentMapper.toDocument(track));
        }
        writer.commit();
    }
}

ByteBuffersDirectory keeps the whole index in heap — ideal for a local library rebuilt at process start. Swap in FSDirectory.open(path) if you need persistence across restarts.

Upsert / delete by id

public void upsert(Track track) throws IOException {
    synchronized (writeLock) {
        writer.updateDocument(
                new Term(TrackDocumentMapper.ID, track.id()),
                TrackDocumentMapper.toDocument(track));
        writer.commit();
    }
}

public void deleteById(String trackId) throws IOException {
    synchronized (writeLock) {
        writer.deleteDocuments(new Term(TrackDocumentMapper.ID, trackId));
        writer.commit();
    }
}

Close

writer.close();
directory.close(); // order matters — writer first

Serialize mutations with a lock if the index is shared across request threads.

Keep the index warm and consistent

Rebuild once from the source of truth at startup. Prefer upsert / delete by id for single-row edits; full rebuild for bulk operations or when sync fails. Never treat Lucene as authoritative.

Real numbers (~4k tracks)

On a local music library of 4 322 tracks (in-memory ByteBuffersDirectory), a full rebuild looks like this:

MetricValue
Documents4 322
Wall time~400 ms
Memory used~1.1 MB

So for a few thousand beans, a full rebuild is cheap enough to run at startup — and even as a fallback when incremental sync fails. (A suggest dictionary, if you add one later, sits in its own Directory and adds a little more RAM.)

The inverted index

Term bob on field title: posting list of docs that contain that token.

Term bob on field title: posting list of docs that contain that token.

After commit, Lucene does not keep a bag of words on each document. It keeps an inverted map: term → documents (the posting list). Type bob on title and you read every track whose title tokenized to bob — including Free (Bob Sinclar Remix).

Same idea on artist. Six tracks, four distinct names after analysis:

DocsArtist (stored)
1, 3Bob Sinclar
2Bob Marley
4, 5Claude François
6François Valery

Lucene does not store that table. It stores the sorted inverted map — lowercased, ASCII-folded (François → francois):

artist:bob       →  1, 2, 3
artist:claude    →  4, 5
artist:francois  →  4, 5, 6
artist:marley    →  2
artist:sinclar   →  1, 3
artist:valery    →  6

That posting list is what you use when searching. We will talk about this in the next article.

The full demo lives on GitHub: lucene-search-tracks.

© 2010 - 2026 David Pilato

Search is powered by Pagefind. Just hit CTRL+K or CMD+K to start searching.

Powered by Hugo with Dream and Devrel themes.

Details

I discovered Elasticsearch project in 2011. After contributed to the project and created open source plugins for it, David joined elastic the company in 2013 where he is Developer and Evangelist. He also created and still actively managing the French spoken language User Group. At elastic, he mainly worked on Elasticsearch source code, specifically on open-source plugins. In his free time, he likes talking about elasticsearch in conferences or in companies (Brown Bag Lunches AKA BBLs). He is also author of FSCrawler project which helps to index your pdf, open office, whatever documents in elasticsearch using Apache Tika behind the scene.

Who am I?

Developer | Evangelist at elastic and creator of the Elastic French User Group. Frequent speaker about all things Elastic, in conferences, for User Groups and in companies with BBL talks. In my free time, I enjoy coding and deejaying as DJ Elky, just for fun. Living with my children in Cergy, France.

Social Links