docs: make the issue tracker the source of truth
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m28s
CI / e2e (pull_request) Successful in 6m9s

Work has been starting from a chat message and a plan file, so two
people could pick up the same thing and neither could see the other.
The tracker is where that is visible.

Search before starting, claim before the first edit -- not before the
commit, since the point is that the other person can see the work is
taken while it is being done. If no issue covers it, open one first:
that is what makes the tracker a description of the project rather than
a description of the past.

The conventions were already right and are written down rather than
reinvented -- the Kind/Area/Priority/Platform/Reviewed/Status taxonomy,
its exclusive scopes, #73 as the roadmap, real Gitea dependencies for
hard blockers, and PR #83's body shape.

What #83 also demonstrated is that a Closes list closes nothing
reliably: it listed ten and five of them sat open in main for a
fortnight. So closing is a step you take and verify, not a keyword you
trust.

.planning/ stops being a queue and keeps design documents and measured
history -- NOTES.md, the audits, the completed plans and the arguments
in them. plans/pending/ is gone, because a plan nobody is executing is
an issue; everything unimplemented in it is now #85-#91, and each
completed plan says which issue carries its remainder. autotag.md is
kept as a historical record, marked stale where the scoring overhaul
overtook it.

The commit grammar is unchanged and is load-bearing for a different
reason, so the issue number lives in the branch name and the PR body
rather than the commit subject.

Refs #92
This commit is contained in:
2026-08-18 16:23:52 -04:00
parent ae82fd2233
commit eb139cf872
8 changed files with 97 additions and 202 deletions
@@ -1,383 +0,0 @@
# 015 — Android release pipeline
Ship an Android APK from CI on every version tag, published to the Gitea
generic package registry so Obtainium can poll a plain URL.
The baseline is `~/Development/ljos`, whose `.gitea/workflows/ci.yml`
`android:` job has been through the failure modes already. Most of what
follows is a transcription of that job onto this repo's conventions;
where it differs, the difference is argued.
## What this is not
**This ships a pipeline, not a usable Android music player.** The
success criterion is a signed, installable APK that launches — not an
app anyone would want. Explicitly out of scope, and each is real:
- `backend/mediacontrols/mpris_linux.go` **will be compiled on Android**.
Go's `android` GOOS implies the `linux` build tag, so the `//go:build
linux` file is in the build and MPRIS will look for a session bus that
does not exist. It compiles; it will error at runtime.
- `backend/system` resolves XDG paths. Android has no XDG.
- The explore catalog artifact is ~0.6 GB. Nothing on a phone wants that.
- The shell is a desktop shell: an eleven-item sidebar, a 800×600
measured minimum, a transport bar. None of that is a phone layout.
- The library scanner walks a filesystem Android does not grant.
Those are the *next* plan, if there is one. Conflating them with this one
is how a build pipeline takes six weeks.
## Phase 0 — the gate [DONE 2026-08-16]
**Passed, further than asked.** No source changes were needed; a full
27 MB fat APK built first try, both ABIs, production-stripped. Numbers,
the environment and four non-obvious findings are in
`.planning/NOTES.md` — including a scaffold bug that put a *debug*
library in the release APK's phone ABI, fixed here.
**It also installs and launches on an emulator, and then exits.** One
line stops it: `backend/system/buildUserDirPath` switches on
`runtime.GOOS` and Android takes the `default:` branch returning
`errUnsupportedOS`, so `main()` hits `os.Exit(1)` six milliseconds
after the JNI bridge comes up. That is the *first* thing that stops it,
not the only one — see the "not this" section above, all of which is
still true and still out of scope.
The emulator tier that found it is now part of the harness:
`scripts/android-emulator.sh`, the `make android-*` targets, and
`.pi/skills/yellowjacket-dev/references/android-tier.md`. It exists
because the failure is invisible in all three places anyone would look
(no panic, no tombstone, no crash buffer) and ActivityManager restarts
the app fast enough that `pidof` always answers — so the tier's
assertion is "same pid after N seconds", not "it started".
Original phase 0 text follows, kept because its reasoning is what the
later phases rest on.
Everything downstream is wasted if the c-shared link fails. Establish it
by hand, locally, before writing a line of YAML.
Already established, by probe rather than by assumption:
```
GOOS=android GOARCH=arm64 CGO_ENABLED=0 go build ./backend/... ./internal/...
```
compiles the entire tree. Exactly two packages fail, and both fail only
because their Android implementation is cgo:
- `ebitengine/oto/v3``driver_android.go` needs the bundled **oboe**
C++ backend. Oto supports Android natively; there is no Java audio
glue to write.
- `wails/v3/pkg/application``mobile_features_android.go` needs the
JNI bridge.
`modernc.org/sqlite` (the whole database layer), `beep`, `godbus` and
every `backend/` package are clean. **No source changes are known to be
required**, which is the single most surprising finding here and the
reason this plan is worth doing at all.
What Phase 0 must actually verify:
1. Install NDK **r26d** (`26.3.11579264`) locally. Pinned, not "whatever
sdkmanager gives you" — ljos's AGENTS.md records newer NDKs breaking
this build.
2. Generate the scaffolding (Phase 1) and run
`wails3 task android:compile:go:shared ARCH=arm64` by hand.
3. Confirm `build/android/app/src/main/jniLibs/arm64-v8a/libwails.so`
exists and is an ARM64 shared object.
4. Repeat for `amd64` (the emulator ABI).
**If the link fails, stop and re-plan.** The likely culprits, in order:
alsa (oto must select oboe, not ALSA — if it reaches for `alsa.pc` the
build tags are wrong), and `main.go`'s `//go:embed all:frontend/dist`
combined with the generated `main_android.gen.go` overlay.
Deliverable: a note in `.planning/NOTES.md` recording the exact command
and the NDK version that produced a `.so`, or the reason it cannot.
## Phase 1 — un-ignore and commit the Android scaffolding [DONE]
Done as a side-effect of phase 0, which could not run without it. One
correction to the text below: **step 1 is wrong.** `update
build-assets` does not generate the android tree (NOTES.md explains);
it was generated with `generate build-assets` into a scratch dir and
`android/` copied across. CLAUDE.md is corrected to match. Steps 2-5
were done as written.
`build/android/` is gitignored (`.gitignore:72`) and its `includes:`
entry was dropped from `Taskfile.yml` during plan 009. That was correct
when nothing could target Android and is what has to be undone.
1. `wails3 task common:update:build-assets` — beta.8 embeds
`internal/commands/build_assets/android/`, so this generates the tree.
2. Remove `build/android/` from `.gitignore`; add `build/ios/`'s reason
to a comment so the asymmetry is explained rather than looking like an
oversight.
3. Add `android: ./build/android/Taskfile.yml` to `Taskfile.yml`'s
`includes:`.
4. **Gitignore the tree's own output**, or the repo grows a few hundred
Gradle intermediates. ljos has exactly this problem — its
`app/build/android/app/build/**` is committed. Ignore:
- `build/android/app/build/`
- `build/android/app/src/main/jniLibs/`
- `build/android/overlay.json` and `build/android/gen/`
5. `make build-prod` and `make test` still pass — the new include must
not perturb the desktop path.
**The refresh hazard has to be written down.** CLAUDE.md's Packaging
section already says `build/`'s platform metadata is regenerated from
`build/config.yml` and hand edits are lost. Phase 2 edits `build.gradle`
by hand. Extend that paragraph to name `build/android/app/build.gradle`
specifically, because the loss is silent and the symptom (a debug-signed
APK) appears months later as a failed update.
## Phase 2 — make the APK identifiable and updatable [DONE 2026-08-16]
**Narrower than planned, because beta.8's scaffold is ahead of ljos's
beta.3: the release signing config already exists** and reads the four
`ANDROID_KEYSTORE_*` variables with a debug-keystore fallback. So this
phase was identity and versioning only. Verified end to end:
| | |
|---|---|
| package | `app.yellowjacket` (was `com.wails.app`) |
| versionCode / versionName | `10301` / `1.3.1`, from `YJ_VERSION_CODE` / `YJ_VERSION` |
| label | `YellowJacket` |
| signing | throwaway keystore -> `Signer #1 DN: CN=YellowJacket Test`, not the debug key |
| ABIs | arm64-v8a + x86_64, both production-stripped |
Installs and launches under the new identity. Still exits on the known
`buildUserDirPath` bug, which is phase 0's finding and not this phase's.
Two things this phase learned that the text below did not know:
- **The identity has to be declared twice.** `applicationId` in
`app/build.gradle` is what Gradle installs; `APP_ID` in
`build/android/Taskfile.yml` is what every adb-driven task targets.
`ANDROID.md` says to set `APP_ID` in `build/config.yml` — that does
nothing in beta.8, verified with `--dry`. Both are set, each
commented pointing at the other.
- **The launcher activity is not under the applicationId.** It stays
`com.wails.app.MainActivity` (the scaffold's Java package), so
`am start -n app.yellowjacket/.MainActivity` resolves the dot against
the wrong package and fails. `scripts/android-emulator.sh` carries the
fully-qualified name and a comment saying why.
The `keytool` PKCS12 note below was confirmed verbatim: given a
`-keypass` differing from `-storepass` it prints "Different store and
key passwords not supported for PKCS12 KeyStores. Ignoring
user-specified -keypass value."
Original phase 2 text follows.
Edit `build/android/app/build.gradle`, following ljos's, whose comments
are worth reading before writing this:
- `applicationId "app.yellowjacket"` — matches `config.yml`'s
`productIdentifier`. The `namespace` stays `com.wails.app` (it is the
Java package, not the app identity).
- `versionCode Integer.parseInt(System.getenv("YJ_VERSION_CODE") ?: "1")`
**`Integer.parseInt`, not `(...) as Integer`**. Groovy binds the
parentheses to `versionCode` first, so the cast reads as
`versionCode("1") as Integer`, which sets a String and then casts the
setter's null return; Gradle fails the whole project with "Value is
null" at that line.
- `versionName System.getenv("YJ_VERSION") ?: "0.0.0"`.
- `abiFilters 'arm64-v8a', 'x86_64'`.
- A `release` signing config reading `ANDROID_KEYSTORE_FILE` /
`_PASSWORD` / `ANDROID_KEY_ALIAS` / `ANDROID_KEY_PASSWORD`, falling
back to the debug keystore only when no keystore is supplied.
**Android orders releases by an integer and refuses anything not greater
than what is installed.** A hardcoded `versionCode 1` means the first
install is the last: every later build is rejected as a downgrade and the
only fix is an uninstall. `1.3.1 -> 10301`, monotonic as long as minor
and patch stay under 100.
**Signing is not optional past the first install.** Android refuses to
update an app whose signing key changed, and the debug keystore differs
between every machine and every runner — so an unsigned CI build is a
decision to reinstall by hand forever. The job must **refuse to build**
without the keystore rather than quietly produce an APK that can never be
updated.
There is **one password and two required secrets**. keytool has defaulted
to PKCS12 since JDK 9 regardless of the `.jks` extension, and PKCS12
cannot hold a separate key password — given `-keypass` it warns and
ignores it. So `ANDROID_KEY_PASSWORD` defaults to the store password and
`ANDROID_KEY_ALIAS` to `yellowjacket`. Asking for a second password that
cannot exist is how someone sets a wrong value and debugs Gradle at
midnight.
Add `make android``PATH="$(TOOLBIN):$$PATH" go tool wails3 task
android:package:fat`, beside `build-prod`. `make skill-check` fails on a
documented target that does not exist, so document it only once it does.
## Phase 3 — the workflow [DONE 2026-08-16]
`.gitea/workflows/android-apk.yml`, plus `docs/android-release.md` as
the operating document its error messages point at (phase 4's
documentation half; the secrets themselves still have to be created by
hand — see the table there).
Three departures from the text below, all argued in the file:
- **No `continue-on-error`.** The plan inherited it from ljos, where
the Android job shares a pipeline with a server deploy that must
never go red over a phone build. Here it is standalone and can
neither delay nor redden anything, so a release step that fails
silently is strictly worse than one that fails visibly.
- **No cached `wails3` binary.** The plan budgeted for ljos's
`tools-bin` copy. Unnecessary: the CLI is a vendored `go tool`, and
the runner already bind-mounts `GOCACHE`/`GOMODCACHE` for every job,
so it is warm from `ci.yml`'s own `make bindings-check`. The GTK and
WebKit *dev* headers are still installed, because `go tool wails3`
links them.
- **A fourth cache volume, `/cache/gradle`.** Not in the plan and worth
~700 MB a run.
Four publish-gates were added and each was checked against a real APK:
both ABIs present, `versionCode` equal to the one derived from the tag,
a non-empty artifact, and **not signed with the debug key** — verified
by pointing the check at a deliberately debug-signed build, which it
refused.
Rehearsed locally with the exact CI invocation
(`make android ANDROID_SDK=... ANDROID_NDK=...`, `YJ_VERSION`,
`YJ_VERSION_CODE`, a throwaway keystore): `app.yellowjacket`,
versionCode 10301, versionName 1.3.1, label YellowJacket, both ABIs,
`Signer #1 DN: CN=YellowJacket`. Not yet run on the runner.
Original phase 3 text follows.
New file: `.gitea/workflows/android-apk.yml`. **Not a job in `ci.yml`.**
`ci.yml` runs on every branch push and is the workflow that gates; the
runner is capacity 1, and a 45-minute Android build in it would put every
push behind an SDK download.
```yaml
on:
push:
tags: ["v*"]
workflow_dispatch:
```
This is where the baseline genuinely diverges. ljos computes its version
in CI (`scripts/next-version.sh`) and gates the Android job on
`needs.release.outputs.version != ''`, with an `always()` whose absence
would silently kill the manual path. **This repo has no release
automation** — tags are pushed by hand and `homebrew-formula.yml` already
keys on `v*`. So there is no `needs:`, no `always()`, and no status
function to get wrong: the tag *is* the version, and a dispatch falls
back to `git describe --tags --abbrev=0`.
Container, matching `ci.yml`'s conventions (`ubuntu:24.04`, clone by hand
with `PACKAGE_TOKEN` rather than `actions/checkout`, which is a JS action
needing node before any step has installed it):
```yaml
container:
image: ubuntu:24.04
volumes:
- /home/logan/docker/gitea/data/runner/cache/tool:/cache/tool
- /home/logan/docker/gitea/data/runner/cache/android-sdk:/cache/android-sdk
```
The SDK path must be inside the runner's `valid_volumes` allowlist —
a directory outside it makes the job **fail to start**, not silently skip
the mount. `/cache/tool` is already allowed and already holds the Go
toolchain `ci.yml` downloads.
`continue-on-error: true` and `timeout-minutes: 45`. Advisory, because a
tag's other three workflows must not go red over a phone build, and a
backstop because a wedged SDK download must not hold the only runner slot
for hours.
Steps:
1. **System packages.** `ci.yml`'s set plus `unzip` and `openjdk-17-jdk`.
`libasound2-dev` stays — it is for the *host* `wails3` build, not the
Android cross-build, which uses oboe.
2. **Go toolchain** — reuse `ci.yml`'s `/cache/tool/go` block verbatim.
3. **Android SDK and NDK (cached).** ljos's `install_if_missing`
idempotent guard, unchanged: cmdline-tools 11076708, `platform-tools`,
`platforms;android-34`, `build-tools;34.0.0`, `ndk;26.3.11579264`.
sdkmanager is itself idempotent but still spends minutes verifying,
which is why the explicit directory guards are there. ~3 GB and most of
the job's wall clock on the first run; a directory listing after.
4. **wails3.** Cheaper here than in ljos, which pins
`go install …/wails3@$version` against `app/go.mod`. This repo vendors
the CLI (`go tool wails3`, `scripts/toolbin/wails3`), so the version is
already pinned by `go.mod` and there is nothing to drift. It still
*links* GTK and WebKit, so cache the built binary in
`/cache/android-sdk/tools-bin` keyed on the wails version — and note
ljos's finding that **caching the binary alone turned a slow job into
a broken one**: `wails3` is dynamically linked, so the runtime
packages are needed even on a cache hit. Here they are already in
step 1.
5. **Frontend + codegen.** `pnpm install --frozen-lockfile && pnpm build`
(pnpm, not ljos's npm), then `make generate`. `main.go` embeds
`frontend/dist`, so nothing Go-side typechecks without it.
6. **Decode the keystore.** Refuse to build if `ANDROID_KEYSTORE_B64` is
unset, with the sentence explaining why (Phase 2). Decide the absolute
path *here* and export it via `$GITHUB_ENV`**`${{ env.HOME }}`
evaluates to an empty string in Gitea's expression context**, which
turned `$HOME/x.jks` into `/x.jks` and surfaced as a missing file
fifty-five seconds into a Gradle run.
7. **Build.** Compute `YJ_VERSION_CODE` from the tag, verify the keystore
opens with `keytool -list` *before* Gradle does (Gradle only notices at
`:app:validateSigningRelease`, a minute in, and reports it as a missing
file), then `make android`.
8. **Verify the signature.** `apksigner verify --print-certs`, and print
the SHA-256 with the note that a change to it breaks every future
update. **Nothing here pipes into `head`**: under `set -o pipefail`,
`head -1` exits early, the producer takes SIGPIPE, and the step fails
with 141 *after* printing a perfectly good APK. Use `find … -print
-quit` and a captured variable.
9. **Publish** to `api/packages/${OWNER}/generic/yellowjacket-android`,
authenticating `--user "${OWNER}:${PACKAGE_TOKEN}"` — the same
credential pair `arch-package.yml` already uses, not ljos's
`REGISTRY_USER`/`REGISTRY_TOKEN`. Two copies: a versioned one for
history and a fixed `latest/yellowjacket.apk` that Obtainium watches.
Gitea refuses to overwrite, so delete `latest` first. The generic
registry is readable **without credentials**, which is what lets
Obtainium poll a plain URL with no token and no public source mirror.
## Phase 4 — secrets and documentation
Secrets to create on the repo (all under Settings → Actions → Secrets):
| Secret | Required | Note |
|---|---|---|
| `ANDROID_KEYSTORE_B64` | yes | `base64 -w0 yellowjacket-release.jks` |
| `ANDROID_KEYSTORE_PASSWORD` | yes | |
| `ANDROID_KEY_ALIAS` | no | defaults to `yellowjacket` |
| `ANDROID_KEY_PASSWORD` | no | defaults to the store password |
| `PACKAGE_TOKEN` | already exists | used by `arch-package.yml` |
Write the keytool command, the Obtainium URL and the signing-key warning
into a docs page — this is the part of ljos's setup that lives in
`docs/clients.md` and is referenced from the workflow's error messages,
so the messages have somewhere to point.
Then extend CLAUDE.md's CI section: it currently says "four workflows,
three of them package and publish; only `ci.yml` gates". That becomes
five, with the same sentence still true.
## Order and stopping points
Phase 0 gates everything. Phases 12 are one commit's worth of work and
are verifiable locally without CI. Phase 3 is the only part that needs a
runner, and its first run will be slow and will probably fail once on
something in the SDK step — budget for that rather than treating it as a
setback.
**Stop after Phase 0 if the c-shared link does not work.** Every later
phase is scaffolding for a build that does not exist, and the honest
outcome is a NOTES.md entry saying which package cannot cross-compile and
what it would take.
@@ -1,337 +0,0 @@
# 015 — Multi-artist credits, navigable
## The problem
A track credited to more than one artist has exactly one navigable
artist in this app, and the others are punctuation.
`audio_files` carries `artist_credit` (the credit as tagged, for
display) and `artist_id` (one artist, for grouping and browsing).
`primaryArtist()` (`backend/library/artistcredit.go:53`) resolves that
one artist by *string-parsing* the credit: it strips a " feat. "
clause, and deliberately does not split on `&`, `x`, `with` or `,`
because those appear inside real artist names. So "Lana Del Rey ft.
Sean Lennon" stores Lana Del Rey and discards Sean Lennon entirely,
and "Alina Baraz & Galimatias" stores one artist whose name is the
whole credit.
### What the measurement says
Measured 2026-08-16 against a real 26,069-file library (19,840 mp3,
6,229 flac; 57 unreadable, m4a/ogg not examined), plus an 80+80
MusicBrainz `inc=artist-credits` sample.
- **13%** of a random sample of the library's recordings have more
than one credited artist in MusicBrainz (10 of 79 resolved).
Extrapolates to ~3,250 of the 24,989 files carrying a recording
MBID.
- **0.86%** of files (224) carry any structured multi-artist signal in
their own tags. mp3 carries **zero** files with multiple
`MUSICBRAINZ_ARTISTID` values across 19,840 files; flac has 87.
- **1,286** files say "feat." in `ARTIST`; **1,159 of them (90%)**
have nothing structured behind it. A sample of 80 such files was
multi-artist in MB **80 of 80 times**.
CLAUDE.md currently justifies plan 013's removal of `artist_credit` /
`artist_credit_artist` with "3 credits of 2,823 listed more than one
artist". That figure measured **our own writer**, not the library:
`cachedLinkArtist` was called exactly once per credit
(`e7748f1^:backend/library/library.go:1842`), so a collaboration could
never have been recorded, and the three were resolution collisions on
shared credit text. Dropping the join table was still correct — it only
ever held one row, so it was pure join cost — but the stated evidence
does not support "multi-artist is rare". Correcting that claim is part
of this plan.
### Why the tags cannot answer it
Deriving the decomposition locally, with no network, works **79% of the
time** (169 of 215 files with a multi-value `ARTISTS` tag: mp3 69/105,
flac 100/110), and the failures are systematic rather than random:
```
ARTIST = '2Pac feat. Snoop Dogg, Nate Dogg, Hussein Fatal & Yaki Kadafi'
ARTISTS = ['2Pac', 'Snoop Doggy Dogg', 'Nate Dogg', 'Fatal', 'Yaki Kadafi']
```
`ARTISTS` holds **canonical** artist names; `ARTIST` holds
**as-credited** names. Locating one inside the other fails on
"Snoop Doggy Dogg" vs "Snoop Dogg", on "Fatal" vs "Hussein Fatal", and
on Unicode (`Michel'le` vs `Michelle`, `K-Ci` vs `KCi` — U+2010, not
a hyphen). That distinction is precisely what a join phrase encodes,
and it is why this cannot be a tag-parsing feature.
Two format details that will mislead anyone re-running the probe:
Picard writes `ARTISTS` **slash-joined into one TXXX frame** on mp3 and
as **true repeated Vorbis keys** on flac, so a probe splitting only on
NUL undercounts mp3 to zero.
## The shape
MusicBrainz models a credit as ordered parts, and the credit *string*
is derived from them — `artist_credit.name` is a cached render, nothing
more. Each participant is `(position, artist, name, join_phrase)`,
where `artist` is the MBID (canonical, what you navigate to) and `name`
is the credited spelling (what you display).
**Join phrases are assembly instructions, not disassembly
instructions.** Rendering is a concatenation, never a search:
```
for each (position, artist_mbid, credited_name, join_phrase):
emit link(credited_name -> artist_mbid)
emit text(join_phrase)
```
The link positions are known **by construction**. This is load-bearing:
if we instead located each `credited_name` inside the stored
`artist_credit` text, we would reintroduce the mismatch above — the
stored string may have come from the tags while the parts come from the
catalog, and those **disagree for ~1 in 3 multi-artist files** (61 of
90 sampled credits rendered exactly equal to the tag string).
Divergences seen: `'Skrillex feat. Swae Lee'` tagged vs
`'Skrillex & Swae Lee'` in MB; `'STRFKR'` vs `'Starfucker'`;
`'Zedd feat. Hayley Williams'` vs `'... of Paramore'`. Either MB was
edited after tagging or Picard versions differ; either way the search
would miss or match the wrong span.
So `audio_files.artist_credit` stops being the source of truth and
becomes the **fallback**, used only where there are no parts.
## Where the data comes from
The catalog carries the decomposition; no user ever makes a
per-recording call. Two sources were ruled out first, both cheaply:
- **The canonical dump — which is what CI already pulls
(`dumpimport.go:84-85`) — does not have it.**
`canonical_musicbrainz_data.csv` gives `artist_mbids` (ordered list)
and `artist_credit_name`, but that last column is the *rendered*
string. Splitting it on CI needs the as-credited names, so CI would
fail exactly the way a local parse does.
- **The JSON dumps do not cover the catalog.**
`json-dumps/recording.tar.xz` is 31 MB / 368 MB uncompressed and
holds **153,691 recordings**, not ~35M. Measured against the test
library's 24,885 recording MBIDs: **0.00% overlap, zero rows**. It is
some other subset and is not usable.
That leaves the core dump, **`mbdump.tar.bz2`** (7.1 GB compressed at
the 20260815 export), from
`https://data.metabrainz.org/pub/musicbrainz/data/fullexport/`. Four
members are needed:
| member | why | approx rows |
| --- | --- | --- |
| `mbdump/artist_credit_name` | `(artist_credit, position, artist, name, join_phrase)` — the payload | ~4M |
| `mbdump/artist` | `id -> gid`, since the above references artist *row ids* | ~2.6M |
| `mbdump/recording` | `gid -> artist_credit`, to key credits by recording MBID | ~35M |
| `mbdump/release_group` | same, for album credits | ~2M |
### Coverage is not a concern
Of 24,885 distinct recording MBIDs in the test library, **24,808
(99.7%)** already have an `explore_index` recording row, measured
against a database at 2,052,200 rows — i.e. shipped-artifact coverage,
not a local build's. The popularity filter does not strand the long
tail here.
## Status
- **Phase 1 — done.** `backend/explore/dumpcredits.go` +
`dumpcreditswrite.go`, wired into `dumpimport.go`'s `run` behind its
own `credits_import_done` marker.
- **Phase 2 — done.** `cmd/indexexport` writes the two tables;
`artifactimport.go` reads them behind `artifactHasCredits()`.
- **Phase 4 — done, and it does not need Phase 3.** `explore.GetCredits`
reads the catalog tables keyed on the *recording* MBID, which both
sides of the app already carry — a catalog row has one and so does a
local file (`library.Track.RecordingMBID`). So one binding serves the
Explore pages and the library's own lists, and all ten artist-link
call sites render credits today without a local table.
- **Phase 3 (`file_artists`) — not started, and now an
offline-resilience task rather than a prerequisite.** The table is
deliberately *not* declared yet: nothing writes or reads it, and a
schema file plus a datamap note describing behaviour that does not
exist is a claim the code cannot back. Its remaining
value is that credits currently vanish when the catalog is absent or
still downloading, which is precisely the `no-index` state
`ShelfPage.State` exists to describe. Materialising into
`file_artists` is what makes a library stand on its own.
**Nothing renders yet in practice**, because no published artifact
carries credit tables — every credit falls back to its single link
until an index build with Phase 1 runs and is exported.
**Column layouts are verified against the real 20260815 export**, not
taken from the schema docs — `artist(id, gid, …)`,
`artist_credit(id, name, artist_count, …)`,
`artist_credit_name(credit, position, artist, name, join_phrase)` and
`recording(id, gid, name, artist_credit, …)` were each read out of the
dump. `release_group` shares `recording`'s first four columns and is
the one layout still taken on trust; `ErrDumpShape` turns a wrong guess
into a loud failure rather than a quietly wrong catalog.
**Still unrun: the ingest against the real 7.1 GB dump.** Everything is
covered by tests over a synthetic tar, which cannot catch a surprise in
the other ~35M rows.
### Phase 1 — Ingest credits on CI
New dump stage in `cmd/indexbuild`, behind the `indexbuild` tag with
the rest of `dumpimport.go`'s stages.
**Constraint from `b98840e`:** `cmd/indexbuild` is built
`CGO_ENABLED=0` in a plain `golang` container and must not reach the
Wails `application` package — `TestIndexToolsDoNotImportWails` walks
`go list -deps -tags indexbuild`. Nothing here should need it, but a
new `ServiceStartup` hook on a package this imports is how it comes
back. Go's `compress/bzip2` is pure Go and decompress-only, which is
all this needs.
**Measured, 20260815 export.** Tar members are **alphabetical**, and
that is favourable: `artist` (435 MB), `artist_credit` (414 MB) and
`artist_credit_name` (237 MB) all fall inside the first ~900 MB
compressed, while `recording` and `release_group` come later. So the
maps are complete before the rows that consume them arrive, and no
recording data is ever buffered.
Pure-Go `compress/bzip2` decompresses at **26 MB/s uncompressed /
8.7 MB/s compressed** (measured on a 250 MB prefix, 3.01x ratio) —
**~13.7 min** for the whole file single-threaded, and less because the
stream can stop after `release_group` rather than reading the
`series`/`tag`/`track`/`url`/`work` tail. The 2 MB/s origin throttle
dominates, as it already does for every other dump here.
Do not, however, *depend* on the ordering: assert it and fall back to
buffering if a future export reorders, rather than silently emitting
nothing.
- `artist` -> `map[int32]uuid16` (~2.6M x ~20 B = ~60 MB)
- `artist_credit_name` -> `map[int32][]creditPart` (~4M x ~40 B =
~200 MB)
- `recording` / `release_group` -> emit `gid -> credit_id` **only for
MBIDs already in `explore_index`** (the kept set is ~1.4M x 16 B =
~22 MB), which is what keeps 35M rows from being held
Peak ~300 MB, one sequential pass.
**Only multi-artist credits are stored.** A single-artist credit is
`(name, "")` and is already fully described by `explore_index`'s
`artist_name` / `artist_mbid`; storing it would triple the table for
nothing. Post-filter after loading, once the row count per credit is
known.
New tables (and `datamap` entries, or `TestCatalogCoversSchema` fails
the build — both are `Cache`, matching `explore_index`):
```
artist_credit_part(credit_id, position, artist_mbid, credited_name, join_phrase)
```
with `explore_index.artist_credit_id` as the link. Credits are
**shared** — an album's twelve tracks by one artist share one credit
row — which is the opposite of 013's local verdict, and correctly so:
1:1 in a local library, genuinely many-to-one at 2M-row catalog scale.
### Phase 2 — Ship them in the artifact
`cmd/indexexport` currently creates exactly two tables in the artifact
(`explore_index`, `artifact_meta`, at `cmd/indexexport/*.go:147,170`),
so this is a structural addition, not a column.
Estimated size: ~13% of 1.4M recordings, deduplicated by shared credit,
at ~2.3 parts each — order 400k rows, ~18 MB uncompressed. Against a
~0.6 GB install that is acceptable; it must be measured rather than
assumed before merge.
`artifactimport.go` must read it **only if present**, on the writer
handle where `core` is attached — the `artifactHasTotals()` /
`artifactStoresText()` pattern (`artifactimport.go:145-175`), one step
up from a column to a table. An artifact published before this exists
is still a perfectly good catalog and must import as one that declines
to answer. Adding this to the importer's SELECT list without the probe
is how every already-published artifact starts failing.
`artifactCatalogColumns` gains `artist_credit_id`; it is kept in sync
with the exporter by `TestArtifactColumnsMatchExporter`.
### Phase 3 — Materialize locally
```
file_artists(audio_file_id, position, artist_id, credited_name, join_phrase)
```
`credited_name` is stored **per row**, not looked up from
`artists.name` — that is the Snoop-Doggy-Dogg distinction, and it is
the whole point.
Filled at scan/import time by joining `audio_files.recording_mbid`
against the catalog. **Materialized rather than resolved live**,
because the catalog is a downloaded artifact that can be absent or
still arriving — that is why `ShelfPage.State` has a `no-index` value —
and a library whose track rows lose their artists when the catalog is
missing is worse than today.
That implies a backfill for the case where the catalog arrives *after*
the library was scanned. It registers with `jobs` (progress, cancel)
like every other long pass, and takes a **distinct kind** from
`index-build`, since `job-controls.ts` keys its "you will discard hours
of downloading" confirmation on that kind.
`artists` gains rows for guests who own no files. **This changes what
the artists grid shows** and is an open question below.
### Phase 4 — Render
`utils/explore-link.ts` gains a credit-rendering entry point taking
ordered parts and returning a `TemplateResult`. Every row and detail
view already renders artist names through it, so they inherit
multi-artist links without individually knowing credits exist — the
property that made centralising it worthwhile.
Its existing fallback philosophy already covers the no-parts case: "a
list where some rows are clickable and others silently are not reads as
a bug, not as a statement about metadata." Where there are no parts
(no recording MBID, or no catalog row — ~4% of the test library) render
today's behaviour: the flat `artist_credit` string with one link to the
primary artist. **Do not split the string there.** There is genuinely
no information to split on, and that is the one place the temptation
returns.
`primaryArtist()` stays exactly as it is. It remains the fallback and
is still what `artist_id` means.
## Open questions
1. **Catalog credit vs tagged credit, when they disagree** (~1 in 3
multi-artist files). Rendering the catalog's decomposition is what
makes names navigable; preserving the file's is what makes the app
reflect the user's files. Leaning toward: render the catalog
decomposition, keep `artist_credit` as the fallback string. Wants a
deliberate decision, not an accident.
2. **Do guest artists appear in the artists grid?** Phase 3 creates
`artists` rows for people who own no files. The grid currently means
"artists in your library" and joins `audio_files`. A guest on one
track is arguably in the library and arguably not. Whichever way,
the ownership question stays "is there a file" — that rule does not
bend.
3. **`release_group` credits** are ingested in the same pass for
nearly nothing, but album-artist rendering is a separate surface.
Ship the data in phase 1, render in a follow-up rather than widening
phase 4.
4. **Our own `tagwriter`** does not write `ARTISTS` or multiple
`MUSICBRAINZ_ARTISTID` frames, so autotagging a folder degrades the
very field this rests on — the same shape as the existing
track-totals note. Out of scope here; worth recording.
## Verification
- Coverage: re-run the library probe and assert `file_artists` is
populated for ~13% of files, not ~0.9%.
- `TestCatalogCoversSchema` / `TestLifetimesMatchSchema` for the new
tables.
- `TestIndexToolsDoNotImportWails` still passes with the new stage.
- An artifact **without** the credits table imports cleanly (the
`artifactHasTotals` regression shape).
- Round-trip: a known multi-artist recording renders each name as a
separate link with the correct join phrases between them.