The search index build and library scan both write to the same
single-connection SQLite DB. The index build runs continuous batch
transactions that can starve the scan's clearLibraryTables call,
causing the scan to silently hang without logging.
Fix: decouple index build from SetContext. The build now starts
AFTER the soft scan completes on startup. For full rescans, the
PreClear hook stops the index build, and PostScan restarts it.
Also made StartBuild/StopBuild safe for multiple calls:
- StartBuild is a no-op if already running
- StopBuild is a no-op if not running (no deadlock on done channel)
- done channel created per-build, not in constructor
Three features wired together:
1. 'In Library' badges on explore search results:
CheckLibraryMBIDs Wails binding batch-checks which search result
MBIDs exist in the local library. Green badges render on matching
artist cards and album cards.
2. Artist images on local artist-details page:
Local artist pages now call GetArtistMBID(name) to resolve the
MBID from tags, then GetArtistImageURL(mbid) to fetch the cached
Wikimedia photo. Falls back to initial-letter avatar.
3. Tier 3 search index uses direct MBIDs from tags:
buildTier3Library now reads artists.mbid column (from audio tags)
for direct MBID matching, falling back to name matching for
untagged artists. Eliminates false matches and catches artists
that name matching misses.
Restructure indexOneArtist to run LB discography fetches and MB
artist image resolution concurrently. They use different rate
limiters (LB: 3 req/s, MB: 1 req/s) so they overlap without
contention.
Per artist, the indexer now runs two parallel pipelines:
LB pipeline: top-release-groups + top-recordings
MB pipeline: url-rels → Wikidata P18 → Wikimedia image fetch
All artist images are pre-cached during the index build instead
of being resolved on-demand during search. Total build time
drops from ~105 min (sequential) to ~63 min (parallel, MB-bound).
SearchIndex now takes ArtistImageProvider as a dependency. The
Service constructor creates artistImg before the index so both
can share it.
Two changes:
1. Artist image disk cache: ArtistImageProvider now fetches the
actual image bytes from Wikimedia Commons and caches them on
disk (~/.local/share/yellowjacket/artist-image-cache/{mbid}.jpg).
Returns base64 data URLs, same pattern as CoverArtProxy.
First lookup: resolve URL via MB/Wikidata + fetch image (~2s).
Subsequent: instant from disk cache.
404s cached as empty files to avoid re-fetching.
2. Top results artist photos: the Top Results section now shows
artist images from the artistImageCache, same as the Artists
section. Also shows englishName in the top card display name.
Artist image resolution was hitting musicbrainz.org with 10
concurrent unthrottled requests per search — enough to trigger
MB's rate limit rejection. Two fixes:
Backend: add dedicated 1 req/s RateLimiter for MB url-rels fetches
in ArtistImageProvider. Each fetch waits on the limiter before
the HTTP call. Results are cached 30 days so repeat lookups are
instant.
Frontend: switch loadArtistImages from concurrent fire-all to
sequential await loop. Each artist image loads one at a time,
images appear progressively as they resolve instead of all
failing from rate limit rejection.
The previous implementation called LookupArtist (rate-limited MB API)
before fetching url-rels, wasting a rate limiter slot. The fetchURL
for rels also bypassed the MB rate limiter, risking 503 rejections.
Rewrite: fetch MB url-rels once (direct HTTP, cached 30 days), parse
both image and wikidata relations from the same response, resolve
Wikimedia thumb URL. No dependency on MusicBrainzClient — just the
Cache for storage and a plain http.Client.
Also cleared 34 stale cached empty results from previous failed
resolution attempts that were blocking image lookup.
Add ArtistImageProvider that resolves artist MBIDs to photo URLs:
1. MB url-rels 'image' type → extract Commons filename → thumb URL
2. MB url-rels 'wikidata' type → Wikidata P18 property → thumb URL
3. No image → falls back to initial-letter avatar
Wikimedia Commons thumb URLs constructed via MD5 hash bucketing
(standard Commons URL scheme). Results cached in explore_cache
with 30-day TTL — subsequent lookups are instant.
Frontend: search results and artist detail page show artist photos
in the circular avatar when available. Images load async and
replace the initial-letter fallback on arrival. Artist detail page
fires the image fetch alongside the other 4 parallel data loads.
Architecture supports adding more sources (fanart.tv, etc.) by
extending the resolve() method's source chain.
Replace per-card GetThumbnail calls (10 round-trips) with a single
GetThumbnails batch call that fetches all visible album thumbnails
in one Wails bridge round-trip.
Backend GetThumbnails accepts []ThumbnailRequest and returns
map[mbid]→dataURL. Each request still checks library → disk cache
→ CAA in order, but the bridge overhead is 1 call instead of 10.
Frontend fires loadThumbnails() once after search results arrive.
Album cards render immediately with CAA URL fallback, then re-render
once the batch resolves with cached/local data URLs.
CoverArtProxy now checks three sources in order:
1. Local library (instant) — matches by album+artist name against
the release_groups/cover_art tables. Albums the user already
owns show their local cover art immediately.
2. Disk cache (instant) — previously fetched CAA thumbnails.
3. Cover Art Archive (network) — fetches and caches to disk.
Library index is built once on first access (sync.Once) from a
single SQL query joining release_groups → cover_art → artists.
Keyed by lowercased 'album\x00artist' for exact name matching.
GetThumbnail now takes (mbid, albumName, artistName) so the proxy
can check the library before falling back to CAA. Frontend passes
the album title and artist credit from the search result.
Add CoverArtProxy that fetches cover art from CAA, caches the image
bytes on disk (~/.local/share/yellowjacket/cover-art-cache/), and
returns base64 data URLs via the GetThumbnail Wails binding.
First load: fetches from CAA (rate-limited), caches to disk.
Subsequent loads: instant from disk cache, no network.
404s: cached as empty files to avoid re-fetching.
Frontend explore-view loads thumbnails async via GetThumbnail()
calls that fire during render. Cached thumbnails appear as data
URLs directly in img src, bypassing the browser's HTTP stack.
Uncached thumbnails fall back to the CAA URL while the proxy
fetches in the background, then re-render with the cached version.
Also stores caa_id and caa_release_mbid in the search index's
extra_json for future direct Internet Archive URL construction.
Phases 2 (3 LB popularity POST calls) and 3 (3 MB discography
browse calls) were adding ~3-6 seconds to every search through
rate-limited API calls. Now they only run as a fallback during
first launch before the search index is built.
Once the index is ready (after Tier 1, <5 seconds from startup):
Phase 0: local FTS5 index query (instant)
Phase 1: MB search (3 concurrent calls, ~1s)
Phase 4: merge index hits (instant)
Phase 5: filter and cap (instant)
Search drops from ~4-7s to ~1s. The index already carries
popularity data and covers discography cross-referencing,
making the live API calls redundant.
Rewrite SearchIndex with tiered background build:
Tier 1 — Sitewide instant (<5s, 12 calls): top artists, recordings,
and release groups across 4 time ranges. Searchable immediately.
Tier 2 — Sitewide full discog (~16min, 2881 calls): top 20 RGs +
top 100 recordings per sitewide artist (~1440 unique artists from
all_time/this_year/this_month/this_week union).
Tier 3 — Library artists (~4min, 664 calls): match local library
artist names against known MBIDs, index their full discographies.
Catches the user's personal taste that sitewide misses.
Tier 4 — Similar artists (~24min, ~4300 calls): fetch similar
artists from LB labs for each library artist, index their
discographies. Fans out into the user's taste neighborhood.
Tier 5 — Organic growth (0 calls): BrowseReleaseGroups now writes
to the search index in a background goroutine. Every artist page
view adds that artist's discography to the index for free.
Other changes:
- indexRGsPerArtist bumped 10→20 (96% vs 88% coverage)
- indexRecsPerArtist bumped 10→100 (track-name searchability)
- indexMinPopularity = 50 (cuts noise from long tails)
- Dedicated 3 req/s rate limiter for indexer
- Labs similar-artists endpoint at labs.api.listenbrainz.org
- Each tier marks index as ready on completion so search improves
progressively during the ~44min total build
Create SearchIndex in NewExploreService, start background build in
SetContext (on app startup). Search() now has 6 phases:
Phase 0: query local FTS5 index (instant, no API calls)
Phase 1: concurrent MB search
Phase 2: LB popularity boost
Phase 3: cross-reference artist discographies
Phase 4: merge index hits (prepend new entries, dedup by MBID)
Phase 5: filter and cap
Index hits for release groups/recordings not already in MB results
are prepended so popular albums surface even when MB search can't
find them. scalePopularity() maps raw listen counts to 0-100 scores
via log scaling for compatibility with the blended score system.
After MB search + popularity reranking, browse the discographies of
the top 3 artists and fuzzy-match the full query against album titles.
Matching albums not already in results are injected at the front.
Fuzzy matching uses substring containment with word-level ratio
(handles 'for you tatsuro' → 'FOR YOU' at 0.667) and word overlap
as fallback. Threshold: 0.4 ratio.
Example: 'for you tatsuro' now finds FOR YOU by 山下達郎 even though
MB text search treats 'for' and 'you' as stop words and never
returns it. The album is found via Yamashita's cached discography.
Drop artists and recordings with blended score < 25 after popularity
reranking. Cap each entity slice to 15 server-side. Request 20 from
MB to allow filtering headroom. Reduces payload size and noise.
After MB search returns text-relevance-scored results, fetch bulk
popularity data from ListenBrainz (POST /1/popularity/{artist,
recording,release-group}) for all result MBIDs. Blend scores:
final = 0.6 * mb_relevance + 0.4 * log10_popularity
Log-scale normalization ensures massive artists don't drown out
everything, but popular results rise above obscure exact matches.
Release groups (no MB score) sort by raw popularity.
Three LB POST calls run concurrently — each hits a different
endpoint. All are rate-limited and cached (24h TTL).
Example: searching 'tatsuro' now ranks Tatsuro Yamashita (2.5M LB
listens, score 97) above 'tatsuro' vocaloid producer (4 listens,
score 64) despite the latter being an exact name match on MB.