Plans 013 and 014, the album page that prompted them, and the smaller fixes they turned up. Changelog, largest first. ## The local library is shaped like files, not like MusicBrainz `audio_files` carries its own tags and points at `albums` and `artists`; `file_genres` is the one real many-to-many. `recordings`, `release_group_recordings`, `artist_credit`, `artist_credit_artist`, `recording_genres`, `release_groups` and `release_to_rg` are gone from the local side, and with them a six-way join in every read, a `MIN(release_group_id)` subquery in eleven queries and a first-credited-artist subquery in nine. Measured on a real 25,966-file library, every many-to-many that model expressed was 1:1 in the data. - Ownership is a file. `GetFilePathsByRecordingMBIDs`, `LibraryMBIDIndex.CheckMBIDs`, `collectLibraryEntities` and `pruneStaleLocalCrossReferences` all join `audio_files`, so the 812 orphaned recordings, 216 release groups and 260 artists that library carried are now structurally impossible. - One projection: every track query selects from the `track_metadata` view, one row type, one mapper. Nine hand-rolled copies had drifted far enough to report different years on different screens. - `library_id = 0` means every library, so each list query exists once instead of scoped and unscoped with a branch at every call site. - No migration chain. `sql/schemas/` is the one description of the shape; `sql/migrations/`, `applyMigrations` and `schema_migrations` are squashed away, along with the drift between them that had sqlc generating against a stale schema. - `database.InsertTestTrack` is the one test seeder; twenty test files had been assembling the old FK chain each in its own order. ## The catalog stores its ids as bytes `explore_index`'s three 36-char MBID columns and its entity-type text are 16 raw bytes and a small integer. The table and its six indexes go 780 MB to 405 MB on a real 2,052,200-row catalog, which is why a fresh install is ~0.6 GB rather than ~1.0 GB. - `backend/explore/mbid.go` is the only place the encoding is known; everything above it speaks dashed strings. - `CHECK(length(mbid) = 16)` makes a stringly write fail at the insert rather than silently returning no rows, since SQLite does not coerce between TEXT and BLOB. - The importer asks the artifact what encoding it carries and converts on the way in, so the artifact already published keeps working and no format bump is needed. - `indexRowColumns`/`scanIndexRow` replace four copies of a 22-column list, and `TestStoredEncodingRoundTrips` sweeps every read path. ## An album page that says how much of the album is yours - One question, asked once: is there a file. `filePaths` is filled by a single batched lookup when the tracklist settles, and the badge, the Play count, the dimmed rows and every menu item read it — replacing four claims of decreasing confidence that could show a green tick on an album whose every action did nothing. - Play, Play 7 of 12, or no play button at all. - `total_tracks` on `explore_index` (~2 bytes over 400,677 release groups) and on `audio_files` from tags that have always carried it: a complete MBID-matched album now makes no catalog call at all, where it used to spend the most expensive request the app makes. - A merged cluster shows the running order the most releases agree on, and the version list marks the release you own rather than standing a synthetic entry in for it. - `AlbumReleasesFailed`: a slow fetch is no longer reported as a failed one by a 12-second timer. - Rows not in the library are dimmed in place (with `aria-disabled`) instead of the owned ones wearing a green tick and a legend. ## Caches and cover art get ceilings - Only the three tiers of a cover are stored; the full-resolution copy nothing rendered was 1,134 MB of a 1.4 GB covers directory. - One artist portrait is downloaded and the rest are remembered as URLs — 4.1 GB of a 5.3 GB cache was candidates no code path reads. - `browsedArtBudget` and `httpCacheBudget` bound what an age cannot: the same install held art for 5,770 artists in a 1,301-artist library. - `OrphanedArtistImagesJob` joined a bare MBID onto a sharded directory, so it deleted the rows that were the only record of the files it left behind. `explore.ArtistImageDir` is that layout's one definition now. ## The autotag queue asks whether there is work `tagging_items` was a row per album folder, not a queue, and no query read the `tag_status` column that held the answer. The four queue queries ask the files, which matters most where it is least visible: `startPrefetch` was scoring every album in a tagged library against MusicBrainz. ## Phantom playlist tracks resolve in place An M3U8 imported before its files leaves phantom rows; they now match by path and fall back to position, keep their place in the playlist when resolved, and pair best-first so two phantoms cannot claim the same file. ## Playing a track plays the list it is in Double-click, and Play on a single row's menu, queue the list as displayed with `startIndex` on that row — the album page and the track list used to queue one track and discard the album around it. A multi-row selection still plays exactly itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AfVYUVExXsx1nSWrXN8mAh
4.5 KiB
014 — The catalog's compact encoding, and the denominator it owed
Status: complete (2026-08-16). The encoding landed first; the
per-release-group total_tracks denominator landed with the album page
that spends it.
Branch: none
Created: 2026-08-16
Depends on: nothing
Related: 013 (the database audit, which measured all of this), 010
(owned albums offline), 001 (ship core index)
The encoding
Measured on the real 2,052,200-row catalog, converting it through the shipped schema (not a projection):
| object | before | after |
|---|---|---|
explore_index |
383 MB | 242 MB |
idx_explore_index_artist_mbid |
131 MB | 65 MB |
UNIQUE(mbid) |
99 MB | 54 MB |
idx_explore_index_entity_pop |
47 MB | 28 MB |
idx_explore_caa_release |
17 MB | 11 MB |
the two LOWER() indexes |
101 MB | 3 MB (013) |
| total | 780 MB | 405 MB |
Every row converted with the CHECK constraints live, which is also a
result: no MBID in a real 2 M-row catalog is malformed.
No format bump, and no rebuilt artifact needed. The importer asks
the artifact what encoding it carries (typeof(mbid)) and converts on
the way in if it is the old text form, so the artifact already
published keeps working and the exporter switches whenever CI next
runs. That is strictly better than the version negotiation this plan
originally proposed.
The silent-failure risk the plan was written around was handled by
making the failure loud instead of by avoiding the change: a CHECK on
the column turns a stringly write into an error at the insert, the
22-column projection became one constant and one scanner instead of
four copies, and TestStoredEncodingRoundTrips sweeps every read path
in the package. It found one real bug on its first run — the artifact
probe was asking the read pool, where the attached artifact does not
exist.
The denominator
total_tracks on explore_index, ~2 bytes across 400,677 release
groups. It makes "do I have all of this" answerable offline for an
album whose files declared no total, which is a great deal of any
untagged library and the one thing GetAlbumCompleteness cannot answer
from tags. 010 rightly rejected shipping whole tracklists — the
per-artist track budget truncates them, and a truncated tracklist is a
confident lie about which tracks exist. A denominator has no such
problem, and the album page spends it as one: the numerator stays
local (distinct track numbers on disk), only the denominator is
borrowed, and only where the tags have none.
Four things about it are load-bearing.
It is counted before the popularity filter. cmd/indexbuild counts
the canonical dump's rows per kept release, which is that release's
track count because the dump carries one row per recording per
canonical release. Counting the kept recordings instead would say
"9" about a twelve-track album whose other three nobody has played —
worse than saying nothing, and the same class of lie as the truncated
tracklist. TestDumpImportEndToEnd has an unplayed track on a fixture
album for exactly this: three tracks in the total, two indexed as
recordings.
Zero means "the catalog does not say", which is the same third state the local answer already has. An album neither side can total wears no ring rather than a wrong one.
Adding a column to the importer's SELECT is how you break every
artifact already published. artifactHasTotals() asks the attached
artifact whether the column exists, the same way and on the same handle
as artifactStoresText(), and selects a literal 0 when it does not.
Verified by forcing the probe true: the older shape then fails with
no such column: total_tracks, which is what a shipped build would
have done to a file nobody can re-cut retroactively.
A test seeder that binds the upsert's parameters by hand is not
"breaking where the app breaks". Three of them did, on the argument
that a schema change should fail the tests in the same place — and what
it actually produced was missing argument with index 25, three files
at a time, for a column none of them cares about. They go through
upsertBatch now, which is the one writer, and keep the property they
wanted: a field written to the wrong column still fails there.
Done when
GetAlbumCompleteness's gap is answerable for a catalog album the library has no tags for, with no network call.- The artifact grows by less than a megabyte (~800 kB at 400,677 release groups).
- An artifact published before the column still imports.