feat(database): shape the library like files, and shrink the catalog
CI / check (push) Successful in 3m7s
CI / e2e (push) Canceled after 1m45s

Plans 013 and 014, the album page that prompted them, and the smaller
fixes they turned up. Changelog, largest first.

## The local library is shaped like files, not like MusicBrainz

`audio_files` carries its own tags and points at `albums` and
`artists`; `file_genres` is the one real many-to-many. `recordings`,
`release_group_recordings`, `artist_credit`, `artist_credit_artist`,
`recording_genres`, `release_groups` and `release_to_rg` are gone from
the local side, and with them a six-way join in every read, a
`MIN(release_group_id)` subquery in eleven queries and a
first-credited-artist subquery in nine. Measured on a real 25,966-file
library, every many-to-many that model expressed was 1:1 in the data.

- Ownership is a file. `GetFilePathsByRecordingMBIDs`,
  `LibraryMBIDIndex.CheckMBIDs`, `collectLibraryEntities` and
  `pruneStaleLocalCrossReferences` all join `audio_files`, so the 812
  orphaned recordings, 216 release groups and 260 artists that library
  carried are now structurally impossible.
- One projection: every track query selects from the `track_metadata`
  view, one row type, one mapper. Nine hand-rolled copies had drifted
  far enough to report different years on different screens.
- `library_id = 0` means every library, so each list query exists once
  instead of scoped and unscoped with a branch at every call site.
- No migration chain. `sql/schemas/` is the one description of the
  shape; `sql/migrations/`, `applyMigrations` and `schema_migrations`
  are squashed away, along with the drift between them that had sqlc
  generating against a stale schema.
- `database.InsertTestTrack` is the one test seeder; twenty test files
  had been assembling the old FK chain each in its own order.

## The catalog stores its ids as bytes

`explore_index`'s three 36-char MBID columns and its entity-type text
are 16 raw bytes and a small integer. The table and its six indexes go
780 MB to 405 MB on a real 2,052,200-row catalog, which is why a fresh
install is ~0.6 GB rather than ~1.0 GB.

- `backend/explore/mbid.go` is the only place the encoding is known;
  everything above it speaks dashed strings.
- `CHECK(length(mbid) = 16)` makes a stringly write fail at the insert
  rather than silently returning no rows, since SQLite does not coerce
  between TEXT and BLOB.
- The importer asks the artifact what encoding it carries and converts
  on the way in, so the artifact already published keeps working and no
  format bump is needed.
- `indexRowColumns`/`scanIndexRow` replace four copies of a 22-column
  list, and `TestStoredEncodingRoundTrips` sweeps every read path.

## An album page that says how much of the album is yours

- One question, asked once: is there a file. `filePaths` is filled by a
  single batched lookup when the tracklist settles, and the badge, the
  Play count, the dimmed rows and every menu item read it — replacing
  four claims of decreasing confidence that could show a green tick on
  an album whose every action did nothing.
- Play, Play 7 of 12, or no play button at all.
- `total_tracks` on `explore_index` (~2 bytes over 400,677 release
  groups) and on `audio_files` from tags that have always carried it:
  a complete MBID-matched album now makes no catalog call at all, where
  it used to spend the most expensive request the app makes.
- A merged cluster shows the running order the most releases agree on,
  and the version list marks the release you own rather than standing a
  synthetic entry in for it.
- `AlbumReleasesFailed`: a slow fetch is no longer reported as a failed
  one by a 12-second timer.
- Rows not in the library are dimmed in place (with `aria-disabled`)
  instead of the owned ones wearing a green tick and a legend.

## Caches and cover art get ceilings

- Only the three tiers of a cover are stored; the full-resolution copy
  nothing rendered was 1,134 MB of a 1.4 GB covers directory.
- One artist portrait is downloaded and the rest are remembered as
  URLs — 4.1 GB of a 5.3 GB cache was candidates no code path reads.
- `browsedArtBudget` and `httpCacheBudget` bound what an age cannot:
  the same install held art for 5,770 artists in a 1,301-artist
  library.
- `OrphanedArtistImagesJob` joined a bare MBID onto a sharded
  directory, so it deleted the rows that were the only record of the
  files it left behind. `explore.ArtistImageDir` is that layout's one
  definition now.

## The autotag queue asks whether there is work

`tagging_items` was a row per album folder, not a queue, and no query
read the `tag_status` column that held the answer. The four queue
queries ask the files, which matters most where it is least visible:
`startPrefetch` was scoring every album in a tagged library against
MusicBrainz.

## Phantom playlist tracks resolve in place

An M3U8 imported before its files leaves phantom rows; they now match
by path and fall back to position, keep their place in the playlist
when resolved, and pair best-first so two phantoms cannot claim the
same file.

## Playing a track plays the list it is in

Double-click, and Play on a single row's menu, queue the list as
displayed with `startIndex` on that row — the album page and the track
list used to queue one track and discard the album around it. A
multi-row selection still plays exactly itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AfVYUVExXsx1nSWrXN8mAh
This commit is contained in:
2026-08-16 13:58:15 -04:00
co-authored by Claude Opus 5
parent 1128881e8d
commit e7748f1fd5
208 changed files with 10944 additions and 12104 deletions
+55 -176
View File
@@ -5,6 +5,8 @@ import (
"database/sql"
"fmt"
"strings"
"yellowjacket/backend/database/sql/sqlcgen"
)
// SearchRow holds a single result from an FTS5 or basename search.
@@ -183,203 +185,80 @@ func (d *DB) RebuildSearchIndex() error {
return nil
}
// SearchTrackRow holds a full track result from an FTS5 search,
// matching all 16 columns returned by GetAllTracksWithFullMetadata.
type SearchTrackRow struct {
FilePath string
LengthMilliseconds int64
Title string
ArtistName string
TrackNumber sql.NullInt64
DiscNumber sql.NullInt64
Album string
Genre string
Year int64
Composer string
FileType string
SampleRate int64
BitDepth int64
Channels int64
Bitrate int64
FileSize int64
// trackMetadataColumns is the column list of the track_metadata view,
// in the order sqlc generates TrackMetadatum's fields. The FTS
// searches below cannot be sqlc queries (MATCH is not in its grammar),
// so this is the one place the view's shape is written out by hand.
const trackMetadataColumns = `
tm.id, tm.file_path, tm.length_milliseconds, tm.title, tm.artist_name,
tm.track_number, tm.disc_number, tm.album, tm.genre, tm.year,
tm.release_year, tm.composer, tm.file_type, tm.sample_rate,
tm.bit_depth, tm.channels, tm.bitrate, tm.file_size, tm.library_id,
tm.play_count, tm.last_played, tm.cover_art_path, tm.artist_mbid,
tm.release_group_mbid, tm.recording_mbid, tm.album_id, tm.artist_id`
// scanTrackMetadata reads track_metadata rows into the generated row
// type, so an FTS hit and an ordinary query produce the same Track.
func scanTrackMetadata(rows *sql.Rows) ([]sqlcgen.TrackMetadatum, error) {
var out []sqlcgen.TrackMetadatum
for rows.Next() {
var r sqlcgen.TrackMetadatum
if err := rows.Scan(
&r.ID, &r.FilePath, &r.LengthMilliseconds, &r.Title, &r.ArtistName,
&r.TrackNumber, &r.DiscNumber, &r.Album, &r.Genre, &r.Year,
&r.ReleaseYear, &r.Composer, &r.FileType, &r.SampleRate,
&r.BitDepth, &r.Channels, &r.Bitrate, &r.FileSize, &r.LibraryID,
&r.PlayCount, &r.LastPlayed, &r.CoverArtPath, &r.ArtistMbid,
&r.ReleaseGroupMbid, &r.RecordingMbid, &r.AlbumID, &r.ArtistID,
); err != nil {
return nil, fmt.Errorf("scan track metadata: %w", err)
}
out = append(out, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("iterate track metadata: %w", err)
}
return out, nil
}
// SearchFTSTracks performs a full-text search and returns full track
// metadata for each match. Unlike SearchFTS (which returns only 5
// columns), this includes all 16 fields needed for library.Track.
// SearchFTSTracks performs a full-text search and returns whole tracks.
//
// A library id of 0 means every library. There were two of these, one
// per case, each with its own copy of a sixteen-column projection that
// silently dropped the MBIDs and the play count - which is why the
// caller used to pass zeros for them.
func (d *DB) SearchFTSTracks(
query string, limit int,
) ([]SearchTrackRow, error) {
query string, libraryID int64, limit int,
) ([]sqlcgen.TrackMetadatum, error) {
query = strings.TrimSpace(query)
if query == "" {
return nil, nil
}
ftsQuery := buildFTSQuery(query)
// SAFETY: FTS5 MATCH syntax unsupported by sqlc. Query is parameterized; no string interpolation.
rows, err := d.db.QueryContext(d.Ctx, `
SELECT
tm.file_path,
tm.length_milliseconds,
tm.title,
tm.artist_name,
tm.track_number,
tm.disc_number,
tm.album,
tm.genre,
tm.year,
tm.composer,
tm.file_type,
tm.sample_rate,
tm.bit_depth,
tm.channels,
tm.bitrate,
tm.file_size
rows, err := d.reader().QueryContext(d.Ctx, `
SELECT`+trackMetadataColumns+`
FROM search_index si
JOIN track_metadata tm ON tm.id = si.rowid
WHERE search_index MATCH ?
AND (? = 0 OR tm.library_id = ?)
ORDER BY rank
LIMIT ?
`, ftsQuery, limit)
`, buildFTSQuery(query), libraryID, libraryID, limit)
if err != nil {
return nil, fmt.Errorf(
"FTS track search failed: %w", err,
)
return nil, fmt.Errorf("FTS track search failed: %w", err)
}
defer func() { _ = rows.Close() }()
var results []SearchTrackRow
for rows.Next() {
var r SearchTrackRow
if err := rows.Scan(
&r.FilePath,
&r.LengthMilliseconds,
&r.Title,
&r.ArtistName,
&r.TrackNumber,
&r.DiscNumber,
&r.Album,
&r.Genre,
&r.Year,
&r.Composer,
&r.FileType,
&r.SampleRate,
&r.BitDepth,
&r.Channels,
&r.Bitrate,
&r.FileSize,
); err != nil {
return nil, fmt.Errorf(
"could not scan search track row: %w",
err,
)
}
results = append(results, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf(
"search track row iteration error: %w",
err,
)
}
return results, nil
return scanTrackMetadata(rows)
}
// SearchFTSTracksByLibrary performs a full-text search scoped to a
// specific library and returns full track metadata for each match.
func (d *DB) SearchFTSTracksByLibrary(
query string, limit int, libraryID int64,
) ([]SearchTrackRow, error) {
query = strings.TrimSpace(query)
if query == "" {
return nil, nil
}
ftsQuery := buildFTSQuery(query)
// SAFETY: FTS5 MATCH syntax unsupported by sqlc. Query is parameterized; no string interpolation.
rows, err := d.db.QueryContext(d.Ctx, `
SELECT
tm.file_path,
tm.length_milliseconds,
tm.title,
tm.artist_name,
tm.track_number,
tm.disc_number,
tm.album,
tm.genre,
tm.year,
tm.composer,
tm.file_type,
tm.sample_rate,
tm.bit_depth,
tm.channels,
tm.bitrate,
tm.file_size
FROM search_index si
JOIN track_metadata tm ON tm.id = si.rowid
WHERE search_index MATCH ? AND tm.library_id = ?
ORDER BY rank
LIMIT ?
`, ftsQuery, libraryID, limit)
if err != nil {
return nil, fmt.Errorf(
"FTS library track search failed: %w", err,
)
}
defer func() { _ = rows.Close() }()
var results []SearchTrackRow
for rows.Next() {
var r SearchTrackRow
if err := rows.Scan(
&r.FilePath,
&r.LengthMilliseconds,
&r.Title,
&r.ArtistName,
&r.TrackNumber,
&r.DiscNumber,
&r.Album,
&r.Genre,
&r.Year,
&r.Composer,
&r.FileType,
&r.SampleRate,
&r.BitDepth,
&r.Channels,
&r.Bitrate,
&r.FileSize,
); err != nil {
return nil, fmt.Errorf(
"could not scan library search track row: %w",
err,
)
}
results = append(results, r)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf(
"library search track row iteration error: %w",
err,
)
}
return results, nil
}
// scanSearchRows reads all rows from a query result into a slice.
func scanSearchRows(
rows interface {
Next() bool