The post-scan backfills share MusicBrainz's rate limiters with every page the user can open, and both were FIFO — so a thousand-artist enrichment put an album page behind an hour of queued work. WithBackgroundLane/WithBackgroundPriority add a slower second lane: a marked wait takes no token while any interactive wait is outstanding. It is a context marker rather than a parameter because a backfill calls the same client methods a detail page does. A long backfill also has to be visible and stoppable, so jobs.KindCatalogEnrich registers both with progress and cancel — after the work is counted, since these passes are a no-op on every launch once the library is covered. What it does not fetch is the point. It ran for hours against a 900-artist library and marked nothing, because three of the four things it did per artist were work nobody asked for: similar artists, which the artist page already resolves on view, and a full GetArtistImage (fanart.tv, TheAudioDB, Wikidata, Wikipedia, ten portraits) reached only to warm the MB artist lookup EnsureArtistRels does alone. It was also serial across artists while every limiter is per-host and idle. The marks are a table rather than more explore_index columns, because artifactimport merges by column list and a flag added there is a second place to remember. BrowseReleaseGroupsAll pages to exhaustion, where the old call silently cut a prolific artist at 100 release groups. One portrait is downloaded now; the rest are remembered as URLs. resolveAllSources downloaded every candidate, up to ten, full size, while nothing reads anything but primary.jpg — 5.3 GB measured on a real cache, 4.1 GB of it unreachable. OrphanedArtistImagesJob is why that survived: it joined the bare MBID onto the images directory, but artist directories are sharded under a two-character prefix, so it named a path that never existed and deleted the rows that were the only record of the files it left behind. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UDCbcCZQepnpSQYJ6SxxZm
32 lines
1.4 KiB
SQL
32 lines
1.4 KiB
SQL
-- Per-artist record of which catalog enrichment passes have completed,
|
|
-- so the owned-artist backfill knows what is left to do.
|
|
--
|
|
-- It is a table rather than more flag columns on `explore_index` for one
|
|
-- reason: the downloaded catalog artifact is merged into that table by
|
|
-- column list (see artifactimport.go), so a flag added there is a second
|
|
-- place to remember, and forgetting it silently wipes every mark on the
|
|
-- next catalog update. These marks are about *this install's* fetching,
|
|
-- which the artifact knows nothing about.
|
|
--
|
|
-- Each column is a separate fetch with its own failure mode, which is
|
|
-- why they are not one boolean: an MB browse failing must not claim the
|
|
-- similar-artists fetch, or vice versa. NULL means "not done" — the
|
|
-- timestamp is for debugging and for any future re-fetch policy, not
|
|
-- for expiry. Nothing expires these today.
|
|
--
|
|
-- `explore_index.discog_fetched` is deliberately NOT duplicated here: it
|
|
-- means "this artist's top release groups and recordings are present",
|
|
-- which the artifact legitimately answers for artists it covers.
|
|
|
|
CREATE TABLE IF NOT EXISTS artist_enrichment (
|
|
artist_mbid TEXT PRIMARY KEY,
|
|
|
|
-- The full MusicBrainz browse-by-artist landed: every release group,
|
|
-- with primary and secondary types. ListenBrainz's top-release-groups
|
|
-- endpoint gives neither the tail nor the types.
|
|
browsed_at DATETIME,
|
|
|
|
-- similar_artist_map has been filled for this artist.
|
|
similar_at DATETIME
|
|
);
|