Files
yellowjacket/docs/index-cache.md
T
logan c03c0b8ec4
Build & publish Arch package / arch-package (push) Successful in 2m33s
CI / check (push) Successful in 3m14s
CI / e2e (push) Canceled after 3m3s
test(database): the next destructive repair fails a test, not a volume
The fix for the dropped catalog pins one table in one wrong shape, which
is the failure that happened. What cost the rebuild was more general: a
destructive repair added at `database.NewDB` -- the chokepoint every
binary in this project shares -- without asking which binary it runs in.
The next one will have a different name and a different reason.

So `TestNoCacheTableIsRetiredHere` asserts the outcome instead: put every
`datamap` Cache table into a shape the schema has moved past, open the
database the way cmd/indexbuild does, and require all of them to still be
there. Driving it from `datamap.ByKind` is what makes it cover tables
nobody remembered -- flipping the policy back fails on five, including
the two artist-credit tables added the same day, where the existing test
fails on one. It asserts the rows survive too, because SQLite does an
implicit DELETE before a DROP and a repair that recreated the table would
look identical. And it accepts an error from `NewDB`, because that is the
documented trade: loud is recoverable, gone is not.

`scripts/index-cache-snapshot.sh` covers the half no test can reach. The
volume holds the only copy of a catalog that costs hours of someone
else's bandwidth to re-derive. `VACUUM INTO` rather than `cp`, since a
byte copy of a live SQLite file is a corrupt file of plausible size; the
resumable staging directory is skipped; and each snapshot is reopened and
asked for its catalog row count before anything is rotated out. A corrupt
source and an empty catalog were both exercised: each exits non-zero,
removes its own output, and leaves the previous snapshots alone.

docs/index-cache.md is the restore, and the reason to bother: a restored
snapshot resolves to `refresh` and folds in the listens since, which is
minutes against the 3-23h this rebuild has been estimating.
2026-08-17 13:56:08 -04:00

4.3 KiB

The index cache, and why it has a snapshot

/srv/yellowjacket/index-cache on the Gitea host is the YJ_HOME the search-index job keeps between runs — .gitea/workflows/index-artifact.yml mounts it at /cache. It holds the catalog every user eventually downloads, and it is the one database in this project that is derived rather than downloaded.

That is the whole reason this document exists. An install with a broken catalog re-fetches the ~0.6 GB artifact and is fine in a minute. This database is what that artifact is cut from, so its only route back is re-streaming the MetaBrainz dumps: hours, at a rate that belongs to someone else's server, holding a runner of capacity 1 the entire time.

What happened on 2026-08-17

A schema repair (fix(database): retire a table whose shape the schema moved past) dropped every table whose live shape disagreed with the schema, before applySchema. Correct for the app. Applied here it deleted the catalog 19 seconds into the first run:

retiring a table ... table=explore_index
  reason="column entity_type is TEXT, schema declares INTEGER"
index maintenance mode=build reason="no completed import yet"

The mismatch was real and deliberate: this database is kept in the older text encoding, which artifactStoresText and sourceColumns exist to tolerate. So it would have been judged stale on every run.

Two things came out of it. retireStaleCache is now a build tag — false under indexbuild, true in the app — and TestNoCacheTableIsRetiredHere asserts the outcome rather than the mechanism, so the next destructive repair fails a test instead of a production volume. And the volume got the snapshot it should always have had, below.

Taking snapshots

scripts/index-cache-snapshot.sh [SOURCE_HOME] [DEST_DIR] [KEEP]

Defaults: /srv/yellowjacket/index-cache, /srv/yellowjacket/index-snapshots, keep 2. On the Gitea host, daily and away from the Monday 04:00 build:

30 5 * * * /path/to/index-cache-snapshot.sh >> /var/log/yj-index-snapshot.log 2>&1

Three properties worth knowing before trusting it:

  • It uses VACUUM INTO, not cp. The database may be open, and a byte copy of a live SQLite file is a corrupt file of plausible size. VACUUM INTO takes a read lock and writes a consistent, compacted copy; it is safe to run while a build is in progress.
  • It does not copy data/explore-staging. That is a resumable checkpoint of work in flight — large, constantly changing, and a build resumes without it. What cannot be cheaply re-derived is the finished catalog, which is in the database.
  • It verifies before it rotates. Each snapshot is reopened and asked for its catalog row count; a run that produces an unreadable or empty file fails loudly, deletes its own output, and leaves the previous snapshots alone. Both paths are exercised, not assumed.

Restoring

Stop anything that might be using the volume first — the job holds it for the length of a build, and the concurrency group (search-index) means a queued run will start the moment one ends.

cd /srv/yellowjacket
mv index-cache/data/yj.db index-cache/data/yj.db.broken   # keep it until you are sure
cp index-snapshots/yj-index-<stamp>.db index-cache/data/yj.db
chown --reference=index-cache/data/yj.db.broken index-cache/data/yj.db

Then dispatch the workflow with mode=auto. A restored snapshot is older than the dumps, so indexbuild resolves to refresh and folds in the incremental listens since — which is minutes, not hours.

Two notes on what a restore does not need. The staging directory can be deleted; it will be rebuilt if a build is needed. And the published artifact is untouched by any of this: users keep downloading the last good one until a run reports complete=true and changed=true republishes.

The trade this leaves open

With Cache tables no longer retired under indexbuild, a future explore_index column change will fail this job loudly — at applySchema, or at the first query naming the column — rather than silently rebuilding. That is the right default: loud is recoverable and a silent day of downloading is not. It does mean the next schema change touching explore_index needs a deliberate plan for this one database: take a snapshot, apply the change to a copy, or accept a rebuild knowingly.