feat: autotag scoring overhaul, dump-based explore index, and lyrics search

Consolidates in-progress work across autotag, explore, and library:

- autotag: beets/Picard-informed scoring engine — ID-first matching, VA
  handling, recommendation tiers, and a merged distance/rank cascade, with
  an eval harness for regression tracking.
- explore: offline MusicBrainz dump import/incremental refresh replaces the
  legacy tier crawl; index-first local search with fuzzy matching and a
  dedicated ranker; disk-free guards for dump downloads.
- library: artist-credit extraction and matching.
- lyrics: owned-library lyric search (FTS) with LRCLIB backfill.

Also: rewrite README to be user-focused, and migrate upstream to
git.ljones.me/yonlu/yellowjacket.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-24 12:14:20 -04:00
co-authored by Claude Opus 4.8
parent d5140395da
commit 65048401e8
117 changed files with 17033 additions and 4767 deletions
@@ -0,0 +1,15 @@
-- Contentless FTS5 index over recording lyrics, enabling
-- "search by a lyric fragment → find the song". The rowid is
-- recordings.id. Only recordings with non-empty lyrics are indexed.
--
-- content='' means the lyric text itself is NOT stored a second time
-- (it already lives in recordings.lyrics); the index keeps only the
-- tokenised inverted index, so it stays compact even for large
-- libraries. contentless_delete=1 lets us delete/reinsert a single
-- row when a track's lyrics change (scan update or LRCLIB backfill).
CREATE VIRTUAL TABLE IF NOT EXISTS lyrics_index USING fts5(
lyrics,
content='',
contentless_delete=1,
tokenize='unicode61 remove_diacritics 2'
);