Vlad Babii

Feature SmithFull stack
idea → spec → ship → review

click ↑ to go back
LinkedIn · opens in a new tab

A scalable, modular document indexing system with cross-module subscriptions

ArchitectureDevelopmentInfrastructureFundamentals

An indexing system for documents that live in many places at once. Small modules, one driver per kind of source, each own their view; they subscribe to each other's changes through a registry, and a shared layer links every copy of a document to one meta-document with merged metadata, labels and search. Every module is a container that finds its peers through the registry, so it scales out per driver and nothing ties it to one host; it could run on Kubernetes as it is. The first system built this way is already running; a second one, for mail and storage, gives AI agents a fast local index to read from.

Same document, many places, no authority

A document can exist in a mailbox, on a storage share and on a website at the same time: three copies, three identifiers, three ideas of what "the document" is. None of them is authoritative. The system doesn't need to know what the documents are about; it only knows which facts about them are current, where each fact came from, and who should be told when one changes.

Drivers, a registry and subscriptions

Adding a source means adding a driver. Nothing else changes.

Each driver talks to one kind of source and owns its view of it completely: its own database, its own pace, its own rate limits. Drivers declare what they can do in a registry; everything else asks the registry "who can do X?" instead of hard-coding a peer, and subscribes to the changes it cares about. Behaviour is described in published contracts, so a driver can be added, replaced, scaled out or switched off without touching any other part.

The running system has 24 services and about 60,000 lines of code. Its drivers cover web sources (one generic driver serves many sites, with each site's extraction code stored as data), mailboxes, browser history, git repositories and e-book files.

  • DriverOne per kind of source; reports a catalog entry for every document and fetches content when asked.
  • RegistryWho provides which capability; consumers discover peers instead of configuring them.
  • Global indexMeta-documents and their primary IDs, merged metadata, labels, the shared vocabulary and search.
  • Change eventsModules subscribe to each other's changes: pushed when something changes, pulled by comparing counters when a push was missed.

Automated subscriptions

A new module joins the cluster and starts receiving data without anyone wiring it up.

When a module that needs data is deployed, it first asks the module authority what exists: which modules provide the capabilities it needs, under which contracts, and which filters they accept.

It then subscribes to each provider with specific filters: the fields it wants, the labels or tags it cares about, and where to start from. Each provider stores the subscription and answers with a cursor.

From there each provider does the work internally. It tags every row that matches the new filter as dirty for that subscriber, so the backfill uses the same path as any later change; it puts the work into its outbox queue with a priority, where repeated updates for the same document collapse; and a sender delivers it with retries.

The new module receives the backfill in batches, then live change events and data, into its own inbox at its own pace. If it restarts, it resumes from its cursor; if it missed a push, it compares counters and pulls whatever is still dirty.

  • Discover through the module authority
  • Subscribe with filters and a cursor
  • Backfill and live use the same path
  • Queued, prioritised delivery
  • Resume from the cursor after a restart
A new module: discover, subscribe, backfill, liveA new module: discover, subscribe, backfill, live
A new module: discover, subscribe, backfill, live
Inside a providing module: match, tag, queue, deliverInside a providing module: match, tag, queue, deliver
Inside a providing module: match, tag, queue, deliver
A subscription's lifecycleA subscription's lifecycle
A subscription's lifecycle

Two tiers: catalog and content

The cost split is the architecture, not a cache.

Tier A, the catalog, is cheap and always synced: title, path, source, size, last-modified, content hash and tags. Search runs on it. A driver can hold a hundred thousand documents in tier A without having fetched a single body.

Tier B, the content, is expensive and fetched on demand: the driver pulls the body, extracts text and structure and reports it to the index. A fetched document can be marked as kept in sync; from then on the driver follows the original and pushes changes. A source that never fetches bodies is still a full member of the index.

  • Catalog for every document
  • Content on demand
  • Kept-in-sync documents
  • Search on the catalog
Drivers, registry and the global indexDrivers, registry and the global index
Drivers, registry and the global index

Meta-documents: one ID, many sources

The same document found in several places becomes one document with several sources.

When a second copy of a document appears, in another source or another driver, the copies are linked as sources of a meta-document, and the meta-document's ID becomes the primary ID everyone else uses. Search results, labels and reading progress attach to the meta-document, not to any one copy.

Its metadata is merged from all sources by the merge policy below. Its content can come from any of them: when someone asks for it, whichever source has it can answer with the part it has or the whole document, so a source that is slow, incomplete or gone doesn't make the document unavailable.

The link between copies is the one thing in the system that can't be rebuilt from the sources, so it is treated as the most valuable data. The same care applies one level down: copies of a long document can number their sections differently, so alignment between sources is verified, by comparing a fingerprint of a few sections, rather than assumed.

  • Meta-document as primary ID
  • Copies linked as sources
  • Metadata merged per field
  • Content from any source, partial or full
  • Verified section alignment

Labels and tags help the indexers

Two sources call the same thing by different names. Tags are resolved against one canonical vocabulary, with types, aliases and parents, when they are indexed; drivers send their raw words and never deal with the mapping.

Labels are added by people: private or shared, on any meta-document. The indexers use both tags and labels to rank and filter search, and a label or an indexed tag can also be an instruction: labelling a document is enough to have its full content fetched and kept in sync.

Many search engines, each built for one job

Adding an indexer is just deploying a service; the data arrives by itself.

There isn't one search engine but several, each optimized for one use case: a Meilisearch index for titles, another for titles plus summaries, and a database full-text engine for filtering by tags and labels. Each one is an ordinary module that subscribes to the documents and fields it needs.

Because migration and subscriptions are part of the architecture, a new indexer costs only its own code: build it, deploy it, and it discovers its providers, subscribes, receives the backfill and then stays live. That runs asynchronously, so a new or rebuilding index never blocks normal operation, and an index that turns out to be wrong can be dropped and rebuilt the same way.

  • Meilisearch: titles
  • Meilisearch: titles + summaries
  • Database full-text: tag and label filters
  • New indexer = build + deploy
  • Backfill and live updates for free
  • Async: never blocks normal operation

A fast tag filter inside the database engine

An optimization in the database-based engine: good enough to ship in hours, and replaceable later.

The database-based search engine needed tag filtering of the form "has X, does not have Y, and rank higher if it also has Z". Each document already had a column listing its labels, so the quickest path was to use it now and improve later: it became the list of all tags and labels as short synthetic words, with a full-text index on it. The query is an ordinary full-text search in boolean mode: required words, excluded words, and optional words that only add to the score.

The words are as short as possible: an underscore, a one-letter type and the ID, like _t42 for a tag or _l7 for a label, so even the shortest one stays above the engine's minimum word length. The list is deduplicated and sorted, so the same set of tags always produces the same column value and an unchanged document is never rewritten. An in-memory map translates between names and tokens in both directions, and because tokens carry IDs, renaming a tag changes nothing in the index.

Better tools exist for this, such as bitmap indexes or a dedicated filter service, but each would have taken much longer to build and run. This one reused an index the database already had, and because every engine sits behind the same search contract, it can be replaced like any other indexer when it stops being enough.

  • has X: +_t42
  • not Y: -_t17
  • Z: optional, adds to the score
  • One column, one full-text index
  • Deduplicated, sorted tokens
  • Renames are free (IDs, not names)

Search plans and saved results

A search can run once, or be saved as a plan that runs again and keeps its results.

A one-off search runs and returns. A saved search becomes a plan, and a plan can have several steps: one step applies a filter, another runs an evaluation function over what is left (a score, a check, a custom rule), and the next step works on that output.

Every run produces a result set that is saved with the plan. It can be reviewed later, used as the input of something else, or served directly as a cached result: asking the same question again returns the stored set instead of searching again.

Because plans are just data, they can be scheduled. Search jobs run during the night, so in the morning the usual searches over millions of documents are already answered and feel instant.

  • One-off or saved as a plan
  • Multi-step: filter, then evaluate
  • Evaluation functions per step
  • Result sets saved for review and reuse
  • Results served as a cache
  • Nightly runs: instant in the morning

Metadata merging is policy, and the policy is data

If you will want to change a value, it shouldn't be a literal in the code.

When several sources report the same document, every field is merged by a strategy stored in a table: max, min, or, union, priority, newest, or a small function for the odd case. Priority doesn't mean "first wins": it fills each empty field by walking the list of sources.

Ranking works the same way. Every rule is a row: a signal, a signed weight and an optional function body. Penalties are as important as bonuses, and a signal without rows simply contributes nothing. Tuning becomes a data change you can diff and revert, and "why did this rank here?" can be answered from the database alone.

  • Per-field merge strategies
  • Rules editable without a deploy
  • Signed weights, penalties included
  • Rules as rows, explainable later

Dirty tracking by timestamps, not queues

Correctness doesn't depend on a message arriving.

Every field has a companion "last changed" timestamp, and every consumer stores the version it last received. A record is dirty when ts > cts. A failed delivery isn't an error to handle; it is a row that is still dirty, and the next sync picks it up.

Subscriptions are symmetric: pushing a change event is the normal path, pulling by comparing counters is the fallback, and the timestamp is the truth. The transport is allowed to lose messages.

The same habits run underneath: every service has an outbox written in the same transaction as its change, and an inbox with priorities (user actions first, bulk work last) where repeated updates for the same document collapse into one. Background jobs reload their settings from the database before every run and start disabled until a person turns them on.

  • dirty = ts > cts
  • Push first, pull to heal
  • Transactional outbox
  • Prioritised inbox with merge keys
  • Jobs start disabled
Dirty tracking: ts > ctsDirty tracking: ts > cts
Dirty tracking: ts > cts

Evolution

Each step added one more kind of plurality, and the design had to absorb it.

It started with one data source to index: a script that fetched and wrote files. A second source meant the same document could now appear twice, so state moved into a database and identity became a question. A third source made it clear that each source had to be isolated in its own module, with its own storage and pace. Then came several search engines and indexers working over the same documents, which is what forced contracts, a registry and subscriptions instead of a shared database.

Version 4 is that last step, built over a couple of months. The same shape is now being reused for mail and storage: the same drivers, registry and meta-documents, with different drivers, so AI agents can search a local index instead of reading every source again.

Stack
Node.js · OpenAPI contracts · MariaDB · Meilisearch · MQTT · Docker · Headless Chromium · Containers (Kubernetes-ready)