Files
yellowjacket/cmd/indexexport/main.go
T
yonluandClaude Opus 5 e7748f1fd5
CI / check (push) Successful in 3m7s
CI / e2e (push) Canceled after 1m45s
feat(database): shape the library like files, and shrink the catalog
Plans 013 and 014, the album page that prompted them, and the smaller
fixes they turned up. Changelog, largest first.

## The local library is shaped like files, not like MusicBrainz

`audio_files` carries its own tags and points at `albums` and
`artists`; `file_genres` is the one real many-to-many. `recordings`,
`release_group_recordings`, `artist_credit`, `artist_credit_artist`,
`recording_genres`, `release_groups` and `release_to_rg` are gone from
the local side, and with them a six-way join in every read, a
`MIN(release_group_id)` subquery in eleven queries and a
first-credited-artist subquery in nine. Measured on a real 25,966-file
library, every many-to-many that model expressed was 1:1 in the data.

- Ownership is a file. `GetFilePathsByRecordingMBIDs`,
  `LibraryMBIDIndex.CheckMBIDs`, `collectLibraryEntities` and
  `pruneStaleLocalCrossReferences` all join `audio_files`, so the 812
  orphaned recordings, 216 release groups and 260 artists that library
  carried are now structurally impossible.
- One projection: every track query selects from the `track_metadata`
  view, one row type, one mapper. Nine hand-rolled copies had drifted
  far enough to report different years on different screens.
- `library_id = 0` means every library, so each list query exists once
  instead of scoped and unscoped with a branch at every call site.
- No migration chain. `sql/schemas/` is the one description of the
  shape; `sql/migrations/`, `applyMigrations` and `schema_migrations`
  are squashed away, along with the drift between them that had sqlc
  generating against a stale schema.
- `database.InsertTestTrack` is the one test seeder; twenty test files
  had been assembling the old FK chain each in its own order.

## The catalog stores its ids as bytes

`explore_index`'s three 36-char MBID columns and its entity-type text
are 16 raw bytes and a small integer. The table and its six indexes go
780 MB to 405 MB on a real 2,052,200-row catalog, which is why a fresh
install is ~0.6 GB rather than ~1.0 GB.

- `backend/explore/mbid.go` is the only place the encoding is known;
  everything above it speaks dashed strings.
- `CHECK(length(mbid) = 16)` makes a stringly write fail at the insert
  rather than silently returning no rows, since SQLite does not coerce
  between TEXT and BLOB.
- The importer asks the artifact what encoding it carries and converts
  on the way in, so the artifact already published keeps working and no
  format bump is needed.
- `indexRowColumns`/`scanIndexRow` replace four copies of a 22-column
  list, and `TestStoredEncodingRoundTrips` sweeps every read path.

## An album page that says how much of the album is yours

- One question, asked once: is there a file. `filePaths` is filled by a
  single batched lookup when the tracklist settles, and the badge, the
  Play count, the dimmed rows and every menu item read it — replacing
  four claims of decreasing confidence that could show a green tick on
  an album whose every action did nothing.
- Play, Play 7 of 12, or no play button at all.
- `total_tracks` on `explore_index` (~2 bytes over 400,677 release
  groups) and on `audio_files` from tags that have always carried it:
  a complete MBID-matched album now makes no catalog call at all, where
  it used to spend the most expensive request the app makes.
- A merged cluster shows the running order the most releases agree on,
  and the version list marks the release you own rather than standing a
  synthetic entry in for it.
- `AlbumReleasesFailed`: a slow fetch is no longer reported as a failed
  one by a 12-second timer.
- Rows not in the library are dimmed in place (with `aria-disabled`)
  instead of the owned ones wearing a green tick and a legend.

## Caches and cover art get ceilings

- Only the three tiers of a cover are stored; the full-resolution copy
  nothing rendered was 1,134 MB of a 1.4 GB covers directory.
- One artist portrait is downloaded and the rest are remembered as
  URLs — 4.1 GB of a 5.3 GB cache was candidates no code path reads.
- `browsedArtBudget` and `httpCacheBudget` bound what an age cannot:
  the same install held art for 5,770 artists in a 1,301-artist
  library.
- `OrphanedArtistImagesJob` joined a bare MBID onto a sharded
  directory, so it deleted the rows that were the only record of the
  files it left behind. `explore.ArtistImageDir` is that layout's one
  definition now.

## The autotag queue asks whether there is work

`tagging_items` was a row per album folder, not a queue, and no query
read the `tag_status` column that held the answer. The four queue
queries ask the files, which matters most where it is least visible:
`startPrefetch` was scoring every album in a tagged library against
MusicBrainz.

## Phantom playlist tracks resolve in place

An M3U8 imported before its files leaves phantom rows; they now match
by path and fall back to position, keep their place in the playlist
when resolved, and pair best-first so two phantoms cannot claim the
same file.

## Playing a track plays the list it is in

Double-click, and Play on a single row's menu, queue the list as
displayed with `startIndex` on that row — the album page and the track
list used to queue one track and discard the album around it. A
multi-row selection still plays exactly itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AfVYUVExXsx1nSWrXN8mAh
2026-08-16 13:58:15 -04:00

333 lines
9.4 KiB
Go

// Command indexexport turns a fully built explore index into the
// compact "core" artifact that ships to users.
//
// The full dump-built index is far too large to distribute (~900MB by
// current budget estimates). The core artifact keeps the most-listened
// artists and their discography slice — enough for Explore to be useful
// on a fresh install — and leaves the long tail to the existing lazy
// per-artist fetch paths.
//
// The artifact deliberately contains no FTS table. The importing client
// inserts these rows into its own explore_index, whose AFTER INSERT
// trigger populates explore_index_fts as a side effect, so shipping a
// search index would be redundant weight.
//
// Usage:
//
// YJ_HOME=/var/cache/yellowjacket-index indexexport -o core-index.db
package main
import (
"database/sql"
"errors"
"flag"
"fmt"
"os"
"path/filepath"
"strconv"
"time"
_ "modernc.org/sqlite"
"yellowjacket/backend/system"
)
// Columns copied into the artifact: the global catalog only.
//
// Deliberately excluded are the per-user columns — in_library,
// is_similar, local_artist_id, local_release_group_id,
// local_recording_id — which describe one person's library and are
// recomputed locally by PopulateLocalCrossReferences after import.
const catalogColumns = `entity_type, mbid, title, artist_name, artist_mbid,
aliases, popularity, listener_count, duration, caa_release_mbid,
release_name, primary_type, secondary_types, release_date, total_tracks,
artist_type, country, disambiguation, sort_name, discog_fetched`
var errNoHome = errors.New(
"YJ_HOME must be set to the directory holding the built index",
)
var errEmptyIndex = errors.New(
"source index has no rows — run indexbuild to completion first",
)
func main() {
out := flag.String("o", "core-index.db", "output artifact path")
artists := flag.Int("artists", 50_000,
"number of top artists (by listen count) to include")
perArtistRGs := flag.Int("rgs-per-artist", 15,
"max release groups per included artist")
perArtistRecs := flag.Int("recs-per-artist", 30,
"max recordings per included artist")
flag.Parse()
if err := run(*out, *artists, *perArtistRGs, *perArtistRecs); err != nil {
fmt.Fprintln(os.Stderr, "indexexport:", err)
os.Exit(1)
}
}
func run(out string, artists, perArtistRGs, perArtistRecs int) error {
if os.Getenv("YJ_HOME") == "" {
return errNoHome
}
dataDir, err := system.GetUserDataDirPath()
if err != nil {
return fmt.Errorf("resolve data dir: %w", err)
}
srcPath := filepath.Join(dataDir, "yj.db")
// Read-only so an export can never disturb a build that is still
// running against the same working directory.
db, err := sql.Open("sqlite",
"file:"+srcPath+"?_pragma=busy_timeout(10000)&mode=ro")
if err != nil {
return fmt.Errorf("open source index: %w", err)
}
defer func() { _ = db.Close() }()
var srcRows int
if err := db.QueryRow(
"SELECT COUNT(*) FROM explore_index",
).Scan(&srcRows); err != nil {
return fmt.Errorf("count source rows: %w", err)
}
if srcRows == 0 {
return errEmptyIndex
}
fmt.Printf("source: %s (%d rows)\n", srcPath, srcRows)
if err := os.Remove(out); err != nil && !os.IsNotExist(err) {
return fmt.Errorf("clear output: %w", err)
}
if _, err := db.Exec(`ATTACH DATABASE ? AS core`, out); err != nil {
return fmt.Errorf("attach output: %w", err)
}
if err := createSchema(db); err != nil {
return err
}
if err := copyRows(db, artists, perArtistRGs, perArtistRecs); err != nil {
return err
}
if err := stampMeta(db, srcRows); err != nil {
return err
}
// DETACH before VACUUM: sqlite cannot vacuum an attached database.
if _, err := db.Exec(`DETACH DATABASE core`); err != nil {
return fmt.Errorf("detach output: %w", err)
}
if err := vacuum(out); err != nil {
return err
}
return report(out)
}
// createSchema builds the artifact's tables. No FTS and no triggers —
// the importing client's own trigger rebuilds its FTS on insert.
func createSchema(db *sql.DB) error {
stmts := []string{
// Column types mirror the app's own explore_index, because the
// export is a straight copy: MBIDs as 16 raw bytes, entity
// types as codes. An importer that meets the older text form
// converts it (artifactSelectColumns), so this is a size change
// and not a compatibility break.
`CREATE TABLE core.explore_index (
entity_type INTEGER NOT NULL,
mbid BLOB NOT NULL,
title TEXT NOT NULL,
artist_name TEXT NOT NULL,
artist_mbid BLOB NOT NULL,
aliases TEXT NOT NULL DEFAULT '',
popularity INTEGER NOT NULL DEFAULT 0,
listener_count INTEGER NOT NULL DEFAULT 0,
duration INTEGER NOT NULL DEFAULT 0,
caa_release_mbid BLOB NOT NULL DEFAULT x'',
release_name TEXT NOT NULL DEFAULT '',
primary_type TEXT NOT NULL DEFAULT '',
secondary_types TEXT NOT NULL DEFAULT '',
release_date TEXT NOT NULL DEFAULT '',
total_tracks INTEGER NOT NULL DEFAULT 0,
artist_type TEXT NOT NULL DEFAULT '',
country TEXT NOT NULL DEFAULT '',
disambiguation TEXT NOT NULL DEFAULT '',
sort_name TEXT NOT NULL DEFAULT '',
discog_fetched INTEGER NOT NULL DEFAULT 0,
PRIMARY KEY (mbid)
) WITHOUT ROWID`,
`CREATE TABLE core.artifact_meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)`,
}
for _, stmt := range stmts {
if _, err := db.Exec(stmt); err != nil {
return fmt.Errorf("create artifact schema: %w", err)
}
}
return nil
}
// copyRows selects the core subset: the top artists by listen count,
// then a bounded slice of each one's release groups and recordings.
//
// The per-artist window mirrors the S2 coverage already in
// dumpcatalog.go — a flat global top-N would give a handful of
// superstars everything and everyone else nothing.
func copyRows(db *sql.DB, artists, perArtistRGs, perArtistRecs int) error {
if _, err := db.Exec(`
CREATE TEMP TABLE core_artists AS
SELECT mbid FROM main.explore_index
WHERE entity_type = 1 /* artist */
ORDER BY popularity DESC
LIMIT ?`, artists,
); err != nil {
return fmt.Errorf("select core artists: %w", err)
}
copied, err := insertSelect(db, `
INSERT INTO core.explore_index (`+catalogColumns+`)
SELECT `+catalogColumns+`
FROM main.explore_index
WHERE entity_type = 1 /* artist */
AND mbid IN (SELECT mbid FROM core_artists)`)
if err != nil {
return err
}
fmt.Printf(" artists: %d\n", copied)
// The entity codes are the catalog's storage form; see
// backend/explore/mbid.go. The exporter copies the local index's
// encoding through unchanged, so the artifact carries it too - and
// the importer accepts either, so an artifact built before this
// still imports.
for _, sel := range []struct {
label string
entity int
limit int
}{
{"release groups", 2 /* release_group */, perArtistRGs},
{"recordings", 3 /* recording */, perArtistRecs},
} {
// The window is over artist_mbid so each artist contributes at
// most `limit` rows, ranked by their own listen counts.
n, err := insertSelect(db, `
INSERT INTO core.explore_index (`+catalogColumns+`)
SELECT `+catalogColumns+` FROM (
SELECT *, ROW_NUMBER() OVER (
PARTITION BY artist_mbid ORDER BY popularity DESC
) AS rn
FROM main.explore_index
WHERE entity_type = ?
AND artist_mbid IN (SELECT mbid FROM core_artists)
) WHERE rn <= ?`, sel.entity, sel.limit)
if err != nil {
return err
}
fmt.Printf(" %-15s %d\n", sel.label+":", n)
}
return nil
}
func insertSelect(db *sql.DB, query string, args ...any) (int64, error) {
res, err := db.Exec(query, args...)
if err != nil {
return 0, fmt.Errorf("copy rows: %w", err)
}
n, err := res.RowsAffected()
if err != nil {
return 0, fmt.Errorf("rows affected: %w", err)
}
return n, nil
}
// stampMeta records what the importing client needs to know: that the
// catalog half is already populated, and which incremental listens
// series the popularity numbers are baselined on, so the incremental
// refresh resumes from the right point instead of reapplying deltas.
func stampMeta(db *sql.DB, srcRows int) error {
series := lookupMeta(db, "listens_applied_series")
built := lookupMeta(db, "dump_import_done")
if built == "" {
built = time.Now().UTC().Format(time.RFC3339)
}
entries := map[string]string{
"artifact_version": "1",
"built_at": built,
"source_rows": strconv.Itoa(srcRows),
"listens_applied_series": series,
}
for k, v := range entries {
if _, err := db.Exec(
`INSERT OR REPLACE INTO core.artifact_meta (key, value) VALUES (?, ?)`,
k, v,
); err != nil {
return fmt.Errorf("stamp %s: %w", k, err)
}
}
return nil
}
func lookupMeta(db *sql.DB, key string) string {
var value string
row := db.QueryRow(
`SELECT value FROM main.explore_index_meta WHERE key = ?`, key)
if err := row.Scan(&value); err != nil {
return ""
}
return value
}
func vacuum(path string) error {
db, err := sql.Open("sqlite", "file:"+path)
if err != nil {
return fmt.Errorf("reopen artifact: %w", err)
}
defer func() { _ = db.Close() }()
if _, err := db.Exec("VACUUM"); err != nil {
return fmt.Errorf("vacuum artifact: %w", err)
}
return nil
}
func report(path string) error {
fi, err := os.Stat(path)
if err != nil {
return fmt.Errorf("stat artifact: %w", err)
}
fmt.Printf("\nartifact: %s (%.1f MB)\n",
path, float64(fi.Size())/(1<<20))
fmt.Println("compress with: zstd -19 -T0", path)
return nil
}