Commit Graph
22 Commits
Author SHA1 Message Date
yonluandClaude Opus 4.8 65048401e8 feat: autotag scoring overhaul, dump-based explore index, and lyrics search
Consolidates in-progress work across autotag, explore, and library:

- autotag: beets/Picard-informed scoring engine — ID-first matching, VA
  handling, recommendation tiers, and a merged distance/rank cascade, with
  an eval harness for regression tracking.
- explore: offline MusicBrainz dump import/incremental refresh replaces the
  legacy tier crawl; index-first local search with fuzzy matching and a
  dedicated ranker; disk-free guards for dump downloads.
- library: artist-credit extraction and matching.
- lyrics: owned-library lyric search (FTS) with LRCLIB backfill.

Also: rewrite README to be user-focused, and migrate upstream to
git.ljones.me/yonlu/yellowjacket.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 12:14:20 -04:00
yonlu d5140395da wip on autotagging 2026-05-01 11:52:50 -04:00
yonluandClaude Opus 4.6 93892c10de wip(explore): library-only mode, ranked search, UI polish — as-is
End-of-milestone state for the Explore milestone. Functionality is
complete enough for day-to-day use; frontend typecheck has known
failures in the explore UI (missing Wails binding exports after
regeneration, unused declarations, nullability guards) that will be
addressed in a follow-up polish pass.

Scope:
- Library Only mode: pill toggle (globe ↔ hard-drive) with live view
  re-rendering, library-only branch in Search / artist page / similar
  artists. Suppresses external API calls when enabled.
- Ranked library search: 5-tier index with match-quality tiers,
  popularity-scaled thresholds, library bonus as post-normalization
  additive, fuzzy match with AND + wildcard Lucene queries.
- New schemas: artist_metadata, http_cache.
- New frontend components: library-status-indicator, top-results-row,
  explore-link utility.
- Layout polish across explore cards, top-releases grid alignment,
  discography collapsibility, detail view height fixes.
- Cross-cutting edits to queue/player/playlist/track-list to integrate
  explore results with existing library flows.

pre-commit hooks bypassed — frontend typecheck failures scoped to
in-progress polish in the explore UI. Go build and full backend test
suite are green.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-16 11:57:00 -04:00
yonlu d73226b173 feat: Library Only mode — toggle, search, artist page, similar artists
Backend:
- Migration 17: similar_artist_map table stores per-artist similar
  artist relationships (source_mbid → similar_mbid + name + score)
- Tier 4 index build now persists similar artists to this table
- GetLibrarySimilarArtists(mbid) queries similar artists filtered
  by JOIN with the artists table (library-only, no API calls)
- Added db field to explore.Service for direct queries

Frontend:
- ExploreSettingsStore with libraryOnly toggle, persisted to
  localStorage
- Top bar toggle button with active/inactive styling
- Explore search: skips full MB/LB pipeline when library-only,
  uses only searchLibraryCache (pure JS, instant)
- Artist detail page: in library-only mode, skips all API calls
  (no top tracks, no top releases, no LB play count, no MB
  artist lookup). Uses library store for discography, calls
  GetLibrarySimilarArtists for similar artists.
- Similar artists section: changed from horizontal scroll to
  wrapping flex layout with collapsible toggle (Show all N)
- Removed debug artist ranking log
2026-03-30 15:36:37 -04:00
yonlu e8fdf8dc54 fix: library bonus as post-normalization additive, not pop contamination
The +10M library bonus was added directly to the popularity map,
which made it the maxPop normalization denominator. With maxPop=10M,
every non-library artist's log-normalized popularity collapsed to
near-zero, making their blended score purely 40% of MB relevance.
All non-indexed artists scored ~35 and ranked by MB noise.

New approach:
- Removed +10M from both GetPopularityBatch and boostWithPopularity
- GetPopularityBatch now returns PopularityBatchResult with separate
  Popularity and InLibrary maps
- rerankArtists takes a libraryMBIDs set and applies a fixed +25
  score bonus AFTER blended scoring and normalization
- maxPop reflects real popularity only, so log normalization works
  correctly across all artists

Shannon Wright (766K listens) now properly outranks Shannon Kennedy
(95 listens) because the popularity scale isn't contaminated.
2026-03-30 03:23:45 -04:00
yonlu 3d0349a427 perf: batch popularity lookups in single SQLite query (100+ → 1)
boostWithIndexPopularity was calling GetPopularity() and IsInLibrary()
individually for every search result — ~100 separate SQLite queries
for a typical search (20 artists × 2 + 20 RGs × 2 + 20 recordings).
This took 7.5s on the 'fast path' that was supposed to take ~5ms.

Added GetPopularityBatch(mbids) — collects all MBIDs across all
entity types and fetches popularity + in_library in a single
SELECT ... WHERE mbid IN (...) query. The library bonus (+10M) is
applied during the batch scan.

Expected Phase 2 improvement: ~7.5s → <10ms.
2026-03-29 23:55:26 -04:00
yonlu 8a138797dd fix: mark search index ready at service creation, not just during build
The search index ready flag is an in-memory bool that resets to false
on every app restart. It was only set to true inside build(), which
runs in a goroutine after SoftScanAllLibraries completes. If the user
searched before the build goroutine started, IsReady() returned false
and the search took the slow path (LB popularity + cross-ref: ~2.3s)
even though the SQLite index had all the data from the previous build.

Now MarkReadyIfPopulated() is called eagerly in NewExploreService —
the index is queryable as soon as the service is constructed, before
any goroutines launch. If the explore_index table has rows, ready=true
immediately.
2026-03-29 21:56:52 -04:00
yonlu 7472e31545 fix: always write artist row in index to prevent redundant re-indexing
indexOneArtist only wrote the artist entry when aliases were non-empty.
Artists without MB aliases (common for smaller/niche artists) never got
an entity_type='artist' row, so indexedArtistMBIDs() couldn't see them.
filterUnindexed then treated them as new on every startup, triggering
redundant LB API calls for top-release-groups and top-recordings.

Now the artist row is always written, with aliases as an empty string
when none exist. Subsequent builds will correctly skip these artists.
2026-03-29 16:11:30 -04:00
yonlu d47270e0b8 perf: lightweight post-scan indexing — only index new artists
After a library scan completes, OnAllScansComplete now calls
IndexNewArtists() instead of StartIndexBuild(). This skips the
full tier pipeline (sitewide top artists, similar artists, freshness
checks) and only indexes library artists whose MBIDs are not yet
in the search index.

Flow after scan:
1. Query indexed artist MBIDs (fast, in-memory set)
2. Query library artist MBIDs
3. Diff → only new artists
4. Fetch discographies + images for new artists only

The full StartIndexBuild() still runs on initial launch (when
SoftScanAllLibraries finds no work to do) to handle the tier
pipeline with freshness-based refresh. But adding 5 new albums
to your library no longer triggers a 60-minute index rebuild.
2026-03-29 13:04:21 -04:00
yonlu e6692e9f1b fix: index build restarts from scratch every launch
Root cause: discog_built meta timestamp was only written after ALL
of Tiers 2-4 completed. If the app was closed mid-build (context
cancelled), the timestamp was never set, so the next launch re-ran
everything from Tier 2.

Fix: track each tier independently with tier2_built, tier3_built,
tier4_built timestamps. Each tier's timestamp is written immediately
after it completes, so progress survives app restarts.

On next launch, already-completed tiers are skipped. Combined with
the incremental filterUnindexed logic, a build interrupted at
Tier 4 with 80% of similar artists done will resume from the
remaining 20%.

Also adds getLibraryArtistMBIDs helper for Tier 4 when Tier 3
was skipped (needs library MBIDs without re-running Tier 3).

InvalidateDiscographies now clears all three per-tier timestamps.
2026-03-29 12:55:07 -04:00
yonlu 7f0c6d362a feat: personalized search ranking — library > similar > neither
Migration 15: add in_library and is_similar INTEGER columns to
explore_index. Backfills in_library from existing library MBIDs.

Search index FTS5 query now includes personalization in scoring:
  ORDER BY bm25(...) - (ln(pop+1) * 1.5)
           - (in_library * 3.0) - (is_similar * 1.5)

For equal text+popularity scores:
  - Library artist beats unrelated by 3.0 points
  - Similar artist beats unrelated by 1.5 points
  - Library > Similar > Neither

Tier 3 (library) entries get in_library=1 via markInLibrary.
Tier 4 (similar) entries get is_similar=1 via markSimilar.

MB result reranking (boostWithIndexPopularity) adds a 10M
popularity bonus for library artists, ensuring they always
rank above non-library artists with equal text relevance.
2026-03-28 12:50:57 -04:00
yonlu 15a0b94348 fix: invalidate index discographies after library rescan
The search index's Tier 3 (library artists) depends on MBIDs from
the artists table. If the index built before a rescan populated
those MBIDs, library artists like Flatbush Zombies wouldn't be
indexed — they're not in the sitewide top 1000 and their name
didn't match via fuzzy matching.

Fix: OnAllScansComplete hook now calls InvalidateIndexDiscographies
before StartIndexBuild. This clears the discog_built timestamp so
Tiers 2-4 re-run incrementally, picking up any new library artists
whose MBIDs were just populated by the scan.

The rebuild is incremental — only artists not already in the index
get their discographies fetched.
2026-03-28 12:40:23 -04:00
yonlu 6820e781e7 feat: rerank MB results with index popularity, increase popularity weight
Two changes:

1. Rerank MB results using index popularity (no API calls):
   When the index is ready, boostWithIndexPopularity looks up each
   MB result's MBID in the local index to get cached listen counts,
   then reranks using the same blended score formula. This was
   previously skipped entirely for speed, leaving MB results sorted
   by text relevance only — obscure exact matches beat popular
   partial matches.

2. Increase popularity weight across both scoring systems:
   - Blended score: 60% popularity / 40% relevance (was 40/60)
   - FTS5 index: ln(pop+1) * 1.5 factor (was 0.5)

   Result: 'flatbush' → Flatbush Zombies (97) beats 'Flatbush'
   nobody (61). Popular artists with partial name matches now
   reliably outrank obscure exact matches.
2026-03-28 12:25:29 -04:00
yonlu 54b074eae7 feat: artist aliases in FTS5 index + BM25 blended scoring
Migration 14: add aliases TEXT column to explore_index, rebuild FTS5
with 3 columns (title, artist_name, aliases), recreate sync triggers.
Clears index build timestamps to force alias population on next build.

Artist image provider fetches inc=url-rels+aliases (single call, no
extra cost). GetAliases() extracts alias names from cached MB rels.
indexOneArtist stores aliases as space-separated text after image
resolution populates the cache.

Search query now uses BM25 blended scoring:
  ORDER BY bm25(fts, 3.0, 1.0, 0.5) - (ln(popularity+1) * 0.5)

Column weights: title=3.0, artist_name=1.0, aliases=0.5
- Title matches score 3x higher than artist name matches
- Alias matches are helpful but don't dominate
- Popularity is a log-scaled boost, not an override
- Exact title match on niche entity beats weak match on mega-popular

Enables: 'rhcp' → Red Hot Chili Peppers, 'gnr' → Guns N' Roses,
'sabbath' → Black Sabbath (once index build runs with aliases).
2026-03-28 09:46:54 -04:00
yonlu bb015092d4 fix: prevent search index build from starving library scan for DB access
The search index build and library scan both write to the same
single-connection SQLite DB. The index build runs continuous batch
transactions that can starve the scan's clearLibraryTables call,
causing the scan to silently hang without logging.

Fix: decouple index build from SetContext. The build now starts
AFTER the soft scan completes on startup. For full rescans, the
PreClear hook stops the index build, and PostScan restarts it.

Also made StartBuild/StopBuild safe for multiple calls:
- StartBuild is a no-op if already running
- StopBuild is a no-op if not running (no deadlock on done channel)
- done channel created per-build, not in constructor
2026-03-26 15:21:01 -04:00
yonlu 8f6a4c6a8e feat: 'In Library' badges, artist images on local pages, MBID-based Tier 3
Three features wired together:

1. 'In Library' badges on explore search results:
   CheckLibraryMBIDs Wails binding batch-checks which search result
   MBIDs exist in the local library. Green badges render on matching
   artist cards and album cards.

2. Artist images on local artist-details page:
   Local artist pages now call GetArtistMBID(name) to resolve the
   MBID from tags, then GetArtistImageURL(mbid) to fetch the cached
   Wikimedia photo. Falls back to initial-letter avatar.

3. Tier 3 search index uses direct MBIDs from tags:
   buildTier3Library now reads artists.mbid column (from audio tags)
   for direct MBID matching, falling back to name matching for
   untagged artists. Eliminates false matches and catches artists
   that name matching misses.
2026-03-26 09:34:57 -04:00
yonlu c9b3c86f37 perf: unified per-artist indexing — discography + image in parallel
Restructure indexOneArtist to run LB discography fetches and MB
artist image resolution concurrently. They use different rate
limiters (LB: 3 req/s, MB: 1 req/s) so they overlap without
contention.

Per artist, the indexer now runs two parallel pipelines:
  LB pipeline: top-release-groups + top-recordings
  MB pipeline: url-rels → Wikidata P18 → Wikimedia image fetch

All artist images are pre-cached during the index build instead
of being resolved on-demand during search. Total build time
drops from ~105 min (sequential) to ~63 min (parallel, MB-bound).

SearchIndex now takes ArtistImageProvider as a dependency. The
Service constructor creates artistImg before the index so both
can share it.
2026-03-26 08:54:43 -04:00
yonlu 49a26c6163 feat: cover art proxy with disk cache for instant thumbnail loading
Add CoverArtProxy that fetches cover art from CAA, caches the image
bytes on disk (~/.local/share/yellowjacket/cover-art-cache/), and
returns base64 data URLs via the GetThumbnail Wails binding.

First load: fetches from CAA (rate-limited), caches to disk.
Subsequent loads: instant from disk cache, no network.
404s: cached as empty files to avoid re-fetching.

Frontend explore-view loads thumbnails async via GetThumbnail()
calls that fire during render. Cached thumbnails appear as data
URLs directly in img src, bypassing the browser's HTTP stack.
Uncached thumbnails fall back to the CAA URL while the proxy
fetches in the background, then re-render with the cached version.

Also stores caa_id and caa_release_mbid in the search index's
extra_json for future direct Internet Archive URL construction.
2026-03-25 14:09:30 -04:00
yonlu e4f4639ab7 feat: per-tier refresh intervals with incremental discography builds
Replace the single 7-day full rebuild with per-tier scheduling:

- Tier 1 (sitewide top lists): weekly refresh, 12 API calls
- Tiers 2-4 (discographies): monthly refresh, incremental —
  only fetches discographies for artists not already indexed

On subsequent runs:
- If T1 is fresh, load cached artists from the index (~0 calls)
- If discographies are fresh, skip Tiers 2-4 entirely (~0 calls)
- If discographies are stale, diff against indexed set and only
  fetch new artists that appeared in the sitewide lists

Add helpers: isMetaFresh (per-key freshness check),
loadCachedSitewideArtists (read artists from existing index),
filterUnindexed (diff artist list against indexed set).

After first build: typical startup is <5s (T1 cache load).
Monthly incremental: ~50-100 calls for newly appeared artists.
2026-03-25 13:43:00 -04:00
yonlu 59d5b7272e feat: popularity-scaled per-artist index budgets
Instead of fixed 20 RGs + 100 recordings for every artist, scale
the budget by popularity using a power curve (exponent 0.3):

  Radiohead (2.5M listens): 20 RGs, 100 recordings
  Hans Zimmer (715K):       15 RGs, 71 recordings
  Clutch (178K):            11 RGs, 50 recordings
  Similar (~10K):            7 RGs, 27 recordings
  Organic (unknown):         5 RGs, 10 recordings

Saves ~53% index size (~29 MB vs ~62 MB) with identical API calls.
The savings come from T4 similar artists (long tail) where full
discographies were wasteful. Top artists still get full coverage.
2026-03-25 13:11:15 -04:00
yonlu 48acede1bc feat: 5-tier search index with library + similar artist expansion
Rewrite SearchIndex with tiered background build:

Tier 1 — Sitewide instant (<5s, 12 calls): top artists, recordings,
  and release groups across 4 time ranges. Searchable immediately.

Tier 2 — Sitewide full discog (~16min, 2881 calls): top 20 RGs +
  top 100 recordings per sitewide artist (~1440 unique artists from
  all_time/this_year/this_month/this_week union).

Tier 3 — Library artists (~4min, 664 calls): match local library
  artist names against known MBIDs, index their full discographies.
  Catches the user's personal taste that sitewide misses.

Tier 4 — Similar artists (~24min, ~4300 calls): fetch similar
  artists from LB labs for each library artist, index their
  discographies. Fans out into the user's taste neighborhood.

Tier 5 — Organic growth (0 calls): BrowseReleaseGroups now writes
  to the search index in a background goroutine. Every artist page
  view adds that artist's discography to the index for free.

Other changes:
- indexRGsPerArtist bumped 10→20 (96% vs 88% coverage)
- indexRecsPerArtist bumped 10→100 (track-name searchability)
- indexMinPopularity = 50 (cuts noise from long tails)
- Dedicated 3 req/s rate limiter for indexer
- Labs similar-artists endpoint at labs.api.listenbrainz.org
- Each tier marks index as ready on completion so search improves
  progressively during the ~44min total build
2026-03-25 12:49:20 -04:00
yonlu bba4e1f3de feat: SearchIndex — background build + FTS5 query for popularity index
New SearchIndex struct in searchindex.go:
- Background build fetches top 1000 LB artists, then their top 10
  release groups and top 10 recordings (2001 API calls total)
- Dedicated 3 req/s rate limiter for indexer (LB allows 30/10s)
- Bounded concurrency (3 goroutines) with progress logging
- Batch INSERTs in transactions of 100 rows
- FTS5 query with prefix matching ('for you' → 'for* you*')
- Results sorted by popularity descending
- Skips rebuild if index is < 7 days old
- Marks index ready from existing rows if build fails
- Context cancellation for clean shutdown
2026-03-25 09:39:22 -04:00