Replace the single 7-day full rebuild with per-tier scheduling:
- Tier 1 (sitewide top lists): weekly refresh, 12 API calls
- Tiers 2-4 (discographies): monthly refresh, incremental —
only fetches discographies for artists not already indexed
On subsequent runs:
- If T1 is fresh, load cached artists from the index (~0 calls)
- If discographies are fresh, skip Tiers 2-4 entirely (~0 calls)
- If discographies are stale, diff against indexed set and only
fetch new artists that appeared in the sitewide lists
Add helpers: isMetaFresh (per-key freshness check),
loadCachedSitewideArtists (read artists from existing index),
filterUnindexed (diff artist list against indexed set).
After first build: typical startup is <5s (T1 cache load).
Monthly incremental: ~50-100 calls for newly appeared artists.
Instead of fixed 20 RGs + 100 recordings for every artist, scale
the budget by popularity using a power curve (exponent 0.3):
Radiohead (2.5M listens): 20 RGs, 100 recordings
Hans Zimmer (715K): 15 RGs, 71 recordings
Clutch (178K): 11 RGs, 50 recordings
Similar (~10K): 7 RGs, 27 recordings
Organic (unknown): 5 RGs, 10 recordings
Saves ~53% index size (~29 MB vs ~62 MB) with identical API calls.
The savings come from T4 similar artists (long tail) where full
discographies were wasteful. Top artists still get full coverage.
Rewrite SearchIndex with tiered background build:
Tier 1 — Sitewide instant (<5s, 12 calls): top artists, recordings,
and release groups across 4 time ranges. Searchable immediately.
Tier 2 — Sitewide full discog (~16min, 2881 calls): top 20 RGs +
top 100 recordings per sitewide artist (~1440 unique artists from
all_time/this_year/this_month/this_week union).
Tier 3 — Library artists (~4min, 664 calls): match local library
artist names against known MBIDs, index their full discographies.
Catches the user's personal taste that sitewide misses.
Tier 4 — Similar artists (~24min, ~4300 calls): fetch similar
artists from LB labs for each library artist, index their
discographies. Fans out into the user's taste neighborhood.
Tier 5 — Organic growth (0 calls): BrowseReleaseGroups now writes
to the search index in a background goroutine. Every artist page
view adds that artist's discography to the index for free.
Other changes:
- indexRGsPerArtist bumped 10→20 (96% vs 88% coverage)
- indexRecsPerArtist bumped 10→100 (track-name searchability)
- indexMinPopularity = 50 (cuts noise from long tails)
- Dedicated 3 req/s rate limiter for indexer
- Labs similar-artists endpoint at labs.api.listenbrainz.org
- Each tier marks index as ready on completion so search improves
progressively during the ~44min total build
New SearchIndex struct in searchindex.go:
- Background build fetches top 1000 LB artists, then their top 10
release groups and top 10 recordings (2001 API calls total)
- Dedicated 3 req/s rate limiter for indexer (LB allows 30/10s)
- Bounded concurrency (3 goroutines) with progress logging
- Batch INSERTs in transactions of 100 rows
- FTS5 query with prefix matching ('for you' → 'for* you*')
- Results sorted by popularity descending
- Skips rebuild if index is < 7 days old
- Marks index ready from existing rows if build fails
- Context cancellation for clean shutdown