feat: autotag scoring overhaul, dump-based explore index, and lyrics search

Consolidates in-progress work across autotag, explore, and library:

- autotag: beets/Picard-informed scoring engine — ID-first matching, VA
  handling, recommendation tiers, and a merged distance/rank cascade, with
  an eval harness for regression tracking.
- explore: offline MusicBrainz dump import/incremental refresh replaces the
  legacy tier crawl; index-first local search with fuzzy matching and a
  dedicated ranker; disk-free guards for dump downloads.
- library: artist-credit extraction and matching.
- lyrics: owned-library lyric search (FTS) with LRCLIB backfill.

Also: rewrite README to be user-focused, and migrate upstream to
git.ljones.me/yonlu/yellowjacket.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-24 12:14:20 -04:00
co-authored by Claude Opus 4.8
parent d5140395da
commit 65048401e8
117 changed files with 17033 additions and 4767 deletions
+41
View File
@@ -19,6 +19,47 @@ function tokenize(s: string): string[] {
return s.match(tokenRe) ?? [];
}
/**
* Loose comparison-only normalization: lowercase, drop punctuation
* (keep letters/digits/spaces), collapse whitespace, trim. Mirrors
* the significant part of the backend's autotag.Normalize() so the
* UI can tell a cosmetic-only difference (case / punctuation /
* whitespace — normalized-equal, score unaffected) from a real one.
* Not exhaustive (no qualifier stripping); it only needs to agree on
* "is this difference purely formatting?".
*/
export function normalizeLoose(s: string): string {
return s
.toLowerCase()
.replace(/[^\p{L}\p{N}\s]/gu, '')
.replace(/\s+/g, ' ')
.trim();
}
/**
* Strict compare-only normalization: like normalizeLoose but also
* drops *all* whitespace, so punctuation that merely changes spacing
* doesn't register as a difference. This is what closes the
* "Rock&Roll" vs "Rock & Roll" gap: the backend's Normalize() deletes
* punctuation without collapsing the surrounding spaces, leaving a
* stray space that scores the pair below 1.0 even though the only
* real difference is punctuation/case.
*/
export function normalizeStrict(s: string): string {
return s.toLowerCase().replace(/[^\p{L}\p{N}]/gu, '');
}
/**
* True when `a` and `b` differ only cosmetically — i.e. by
* capitalization, punctuation, or the spacing punctuation induces.
* Used to decide whether a title change is a real conflict or just
* formatting. Empty `a` (no local value) is never cosmetic.
*/
export function isCosmeticDiff(a: string, b: string): boolean {
if (a === '' || a === b) return false;
return normalizeStrict(a) === normalizeStrict(b);
}
/**
* Compute an inline word/punct-level diff between `a` (old) and
* `b` (new), returning a list of segments suitable for inline