feat: autotag scoring overhaul, dump-based explore index, and lyrics search
Consolidates in-progress work across autotag, explore, and library: - autotag: beets/Picard-informed scoring engine — ID-first matching, VA handling, recommendation tiers, and a merged distance/rank cascade, with an eval harness for regression tracking. - explore: offline MusicBrainz dump import/incremental refresh replaces the legacy tier crawl; index-first local search with fuzzy matching and a dedicated ranker; disk-free guards for dump downloads. - library: artist-credit extraction and matching. - lyrics: owned-library lyric search (FTS) with LRCLIB backfill. Also: rewrite README to be user-focused, and migrate upstream to git.ljones.me/yonlu/yellowjacket. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,15 @@
|
||||
-- Contentless FTS5 index over recording lyrics, enabling
|
||||
-- "search by a lyric fragment → find the song". The rowid is
|
||||
-- recordings.id. Only recordings with non-empty lyrics are indexed.
|
||||
--
|
||||
-- content='' means the lyric text itself is NOT stored a second time
|
||||
-- (it already lives in recordings.lyrics); the index keeps only the
|
||||
-- tokenised inverted index, so it stays compact even for large
|
||||
-- libraries. contentless_delete=1 lets us delete/reinsert a single
|
||||
-- row when a track's lyrics change (scan update or LRCLIB backfill).
|
||||
CREATE VIRTUAL TABLE IF NOT EXISTS lyrics_index USING fts5(
|
||||
lyrics,
|
||||
content='',
|
||||
contentless_delete=1,
|
||||
tokenize='unicode61 remove_diacritics 2'
|
||||
);
|
||||
@@ -3,6 +3,7 @@ CREATE TABLE IF NOT EXISTS playlists (
|
||||
name TEXT NOT NULL,
|
||||
is_smart INTEGER NOT NULL DEFAULT 0,
|
||||
smart_rules TEXT,
|
||||
smart_snapshot_at DATETIME,
|
||||
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
|
||||
);
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
-- tagging_candidates durably stores the scored candidate list for a
|
||||
-- tagging group so it survives process restarts. Without it, only the
|
||||
-- top score (a single REAL on tagging_items) is persisted; the full
|
||||
-- candidate list — releases, tracks, alignments — is recomputed every
|
||||
-- session, re-hitting MusicBrainz whenever the short-TTL http_cache has
|
||||
-- expired. The blob is written once when a group is first scored and
|
||||
-- read back on every subsequent open.
|
||||
--
|
||||
-- candidates holds the JSON-encoded []autotag.Candidate. ON DELETE
|
||||
-- CASCADE ties the blob's lifetime to its tagging_items row: when a
|
||||
-- group's tracks change, the scan path deletes the old group_key row
|
||||
-- (and SQLite, with foreign_keys = ON, drops the stale blob with it).
|
||||
CREATE TABLE IF NOT EXISTS tagging_candidates (
|
||||
group_key TEXT PRIMARY KEY,
|
||||
candidates TEXT NOT NULL,
|
||||
computed_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
FOREIGN KEY(group_key) REFERENCES tagging_items(group_key) ON DELETE CASCADE
|
||||
);
|
||||
Reference in New Issue
Block a user