# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Project YellowJacket is a cross-platform desktop music player built with Go (backend) and TypeScript/Lit (frontend), using the Wails framework to bridge them. It supports MP3, FLAC, OGG Vorbis, and WAV playback. ## Planning Active and historical plans live in `.planning/`: - `.planning/NOTES.md` — gotchas, deferred items, open architecture questions, the "we already considered and rejected" list. - `.planning/plans/active/` — work currently in progress (read first). - `.planning/plans/pending/` — sequenced future work. - `.planning/plans/completed/` — one concise recap per shipped milestone. Numbering is sequential and stable across status moves (a plan keeps its `NNN-` prefix as it migrates between `pending → active → completed`). Abandoned plans are deleted; paused work stays in `pending/`. ## Commands ```bash make dev # Hot-reload development (installs deps, generates code, cleans frontend) make dev-debug # Same as dev but with YJ_LOG_LEVEL=debug make dev-headless # Start headless in the background and return (SEED= to seed) make dev-stop # Stop it (SIGTERM, so shutdown hooks run) make dev-logs # Tail .dev/app.log make testdata # Generate the deterministic fixture music library make bulkdata # Generate the ~50k-track measurement library (BULK_TRACKS=) make sandbox-seed NAME= # Build a seeded YJ_HOME by *running* the app make sandbox-seed-bulk # Same, from the bulk library (minutes; it is a real scan) make perf LABEL= # Measure a running app; writes .dev/perf/.json make perf-compare BEFORE= AFTER= # Print the before/after table make build-dev # Debug build with symbols make build-prod # Production build (stripped, UPX-compressed) make generate # Run code generators (sqlc + templ via go generate) make e2e # Playwright smoke suite against a running dev-headless app make e2e-setup # Install the e2e runner + its browser (once) make ui-test # Vitest component/store suite in a real browser (no app) make ui-visual # Same, including toMatchScreenshot comparisons make ui-setup # Install the Vitest provider's own Chromium (once) make bindings-check # Fail if frontend/wailsjs is stale vs the Go bindings make skill-check # Fail if .pi/ documents a make target that doesn't exist make commit-check # Fail if a commit subject is not a Conventional Commit make lint # golangci-lint v2 (strict), all three build configurations make test # All tests with race detector, all three build configurations make vulncheck # govulncheck for CVEs make setup # Install go tools, frontend deps, git hooks (lefthook) ``` ### Running tests All Go test commands require the `-tags webkit2_41` build tag: ```bash go test -tags webkit2_41 ./... # All tests go test -tags webkit2_41 ./backend/player/ # Single package go test -tags webkit2_41 -run TestName ./backend/player/ # Single test ``` The central index builder is behind a second tag and is **not** covered by the command above — `make test` runs both passes, but a manual run needs it spelled out: ```bash go test -tags "webkit2_41 indexbuild" ./backend/explore/... ./cmd/... ``` `backend/testctl` is behind a third tag and needs its own pass too (`make test` runs all three): ```bash go test -tags "webkit2_41 dev" ./backend/testctl/... ``` Audio playback integration tests require `YELLOWJACKET_INTEGRATION=1`. ### Fixtures and the headless harness `test_data/music_library_test/` is **generated, not committed**: run `make testdata` (~1 s) before anything that needs audio. Tests reach it through `internal/testfixtures`, selecting files by *case* (`CaseCoverDedup`, `CaseUnicode`, `CaseDuplicates`, …) rather than by path, and skip themselves when it has not been generated. The app itself can be run without a blocking window — `make dev-headless` — and driven with `playwright-cli` against the dev server on `:34115`, which is the real app with real bindings on `window.go`, bridged to the same Go backend a desktop window would use. **The operational half of all this lives in the `yellowjacket-dev` skill** (`.pi/skills/yellowjacket-dev/`): which tier to reach for, the exact command sequences, seed lifecycle, and the failure modes worth knowing before you meet them. It is deliberately not repeated here — this section describes what exists, the skill says what to run. Two things ride on top of the headless launch, both from plan 005 phase 3: - **The event bridge.** `.playwright/cli.config.json` loads `.playwright/init-events.js` as an `initScript`, which records every backend event on `window.__yjEvents`. Half this app is push-driven, so assertions **await an event, not a timeout**: `await window.__yjEvents.wait('LibraryScanComplete', {timeoutMs: 60000})`. It also provides `ready()` and `call('queue.Queue.GetState', [])`, which times out instead of hanging. - **The dev-only control surface**, `backend/testctl`, mounted at `/__test/` on the same port: `health`, `db/snapshot`, `db/restore`, `emit` (force any backend event, which renders push-driven views without staging the work that would produce them) and `sql`. It is compiled out of non-dev builds and additionally requires `YJ_TESTCTL=1`, which `dev-headless.sh` sets and `make dev` does not. Frozen regression specs live in `e2e/` (its own npm package, so the Vitest browser mode does not share a package with the Playwright runner): `make e2e` against an already-running app. **A fifth tier exists for questions whose answer is a number**, not a pass: `make bulkdata` generates a ~50 000-track library (11 s, 466 MB, gitignored), `make sandbox-seed-bulk` seeds from it by running the app like any other seed, and `make perf LABEL=x` measures startup, the bundle's shape and each view's first open, keystroke cost, what a finished track provokes, what one favourite toggle costs, what sitting idle on Settings costs, and heap after a scripted browse. A measurement that needs state the seed does not have stages it itself, idempotently, so a before and an after see the same shape — the favourite number is meaningless against the seed's one empty playlist, so it builds ten 500-track ones first. It wraps every bound Go method, so "did that refetch the library" is a fact rather than an inference. It is not a spec and does not run in CI. **The cheapest tier needs none of that.** `make ui-test` runs 672 Vitest tests in a real Chromium in ~2 s with no Wails, no backend, no seeded library and no virtual display, because `frontend/wailsjs/` is a pure passthrough to `window.go` / `window.runtime` and `frontend/test/support/wails-fake.ts` replaces just those two globals — so the tests exercise the real generated bindings and the real store code. **`frontend/wailsjs/` is generated by `wails`, not `go generate`**, so the pre-commit codegen check does not cover it. `make bindings-check` (~1.5 s, also a pre-commit hook) regenerates it and fails on a dirty tree; `make bindings` regenerates it for real. **Seeds are produced by running the app**, never by hand-writing a `config.toml` and DB rows — the same discipline `sql/schemas/` gets, for the same reason. See `.planning/plans/completed/005-agent-development-harness.md`. ## Architecture **Wails app lifecycle** (`main.go` → `backend/app.go`): `YellowJacketApp` is the root struct bound to Wails. Its methods are callable from the frontend. Lifecycle hooks: `OnStartup` (init audio), `OnDomReady` (start library scan), `OnBeforeClose` (save window state), `OnShutdown` (persist player/queue state). **Backend packages** (under `backend/`): - `player` — Audio playback via beep. `BufferedStreamer` provides a ring buffer for smooth seeking. - `queue` — Track queue with shuffle (Fisher-Yates), repeat modes, auto-advance, and session persistence. - `library` — Concurrent library scanning, metadata extraction, cover art deduplication, incremental rescan. - `database` — SQLite via pure-Go driver. Schema in `database/sql/schemas/`, queries in `database/sql/queries/`. **sqlc** generates Go code into `database/sql/sqlcgen/` — never edit that directory by hand. **Schema changes need two things, not one.** `sql/schemas/*.sql` is `CREATE ... IF NOT EXISTS` and is what sqlc reads — it's the single source of truth for "what the schema looks like right now", and it's what a fresh install gets verbatim. But it's a no-op against a database that already has the table, so an existing install needs a matching file in `sql/schemas/../migrations/` (e.g. `NNNN_description.sql`, `ALTER TABLE ... ADD COLUMN ...` / `CREATE INDEX ...`) to actually reach that shape. Both run on every open, migrations after schema files, tracked in `schema_migrations` so each applies once; a migration's `ALTER TABLE ADD COLUMN` failing with "duplicate column name" on an already-current database is expected and tolerated, not an error. A few things that bite if forgotten: - **Column order must match between the two paths.** `ALTER TABLE ADD COLUMN` always appends at the end, so a migrated column must also be declared *last* in the `CREATE TABLE` in `sql/schemas/` — otherwise a fresh install and an upgraded install disagree on column order, and a `SELECT *` query (sqlc binds those positionally) silently reads the wrong field on one of them. See `backend/database/migrations_test.go`'s `TestMigrations_ColumnOrderMatchesFreshInstall`, which is the regression test for exactly this. - **Don't put an index on a migrated column in `sql/schemas/`.** Schema files run before migrations, against a database that may not have that column yet — the index's predicate would fail (this is precisely the bug an earlier session shipped and a user hit at `make sandbox`). Declare it in the migration file instead, after the `ALTER TABLE` that adds the column. - This project **had** a 48-step migration chain before and tore it out (see `.planning/NOTES.md`, "No migration chain") because `sql/schemas/` had drifted from what the migrations actually produced and sqlc silently generated against the stale version. The design here avoids that by keeping `sql/schemas/` as the literal target shape (not a hand-maintained description of it) and letting migrations replay tolerantly against it — but the same drift is possible again if a schema change ships without updating both files. Don't reintroduce a *second* description of the schema anywhere else. - **A write wearing a query's shape still needs the writer.** `DB.QueryContext`/`QueryContextWith`/`QueryRow` route to a *query-only* read pool (a second `sql.DB` over the same file), so an `INSERT ... RETURNING` issued through one fails at runtime with "attempt to write a readonly database (8)" — which is exactly what `CreateSmartPlaylist` did, meaning no smart playlist could be created at all. Use `ExecContext`, or `QueryRowWriter` when the statement really does return a row. Nothing caught this because `NewTestDB` shares one in-memory connection and leaves `readDB` nil, so `reader()` returns the *writer* under test and the unit tests exercised a handle the app does not have. `TestNoWritesOnTheReadPool` walks the tree for it, in the same spirit as `TestNoDirectRuntimeEmits` and for the same reason — a lint pass only sees one build configuration. - **Squashing is fine pre-1.0.** While this hasn't shipped to real users, periodically folding `sql/migrations/` into `sql/schemas/` and deleting the migration files (then wiping your own dev/sandbox DB) is a legitimate way to keep the migrations directory from accumulating dev-only churn — same effect as the old "just nuke it" workflow, opt-in instead of mandatory. Stop doing that once real user databases exist in the wild. - `metadata` — Tag extraction (ID3v2, Vorbis Comments, FLAC). - `jobs` — The registry every long-running operation reports through: progress, pause/cancel, a global indicator and (for scans) a pause that survives a restart. Library scans, index builds, downloads and the autotag apply are registered; anything that is not registered has none of that, which is exactly how the three gaps the audit found came about. - `config` — TOML-based settings. Settings page uses HTMX + templ for server-rendered HTML fragments. - `playlist` / `smartplaylist` — Playlist CRUD and rule-based smart playlists. - `mediacontrols` — MPRIS integration on Linux via D-Bus. - `system` — OS-specific paths (XDG on Linux, `%LOCALAPPDATA%` on Windows). - `explore` — Catalog search and browse over `explore_index`. See below. Its **shelves** (`shelves.go`) are the page Explore shows before anyone types, on `home`'s terms — a shelf is a reason, it carries the sentence that says so, and an empty one is omitted. Queries return `explore_index` row ids and are joined back by `rowsByIDs`, so a card has one definition (`artistFromIndex` / `releaseGroupFromIndex` / `recordingFromIndex`, shared with the search path, which is where they were inlined). Three things about it are load-bearing. **"No shelves" is three different statements here and the page says which** — Home can omit an empty shelf honestly, because a library with no history really has less to say, but Explore's data is a *downloaded artifact* that can be absent or still arriving, so `ShelfPage.State` is `ready`, `building` or `no-index` and the empty page names the missing catalog and points at Settings. **Whether there is a catalog is asked of the database, not of a flag**: `GetIndexStatus().TotalRows` is refreshed only between build tiers (0 beside a full catalog on an ordinary launch) and `IsReady()` is set once at startup (so rows staged by a spec afterwards are invisible) — both are the shape `emitStatus` warns about, and one `SELECT 1 … LIMIT 1` cannot be stale. And **two shelves with disjoint ids still repeat each other**: ordered by raw listen count the top albums are one act and its members and the artists row underneath was the same people, which `home`'s duplicate guard cannot see because the rows hold different entity types. Shelves are one album per artist and skip whoever a row above already showed. Found by reading a screenshot. Two of the four shelves the plan named **cannot be built**, and the schema decides that rather than the design: `explore_index` has no genre column to join a genre shelf to, and `similar_artist_map` is not in the shipped artifact and is filled lazily from the network, so a "similar artists" shelf is empty exactly when the page most needs content. The library-joining shelf reads `in_library`, which is set by MBID, so it is correctly absent on an untagged library — the fixture one included. - `home` — The home page's "start listening" shelves. Each shelf is a *reason* (what you played last, what you never played, a genre you have depth in) rather than a filter, and carries the sentence that says so. Its queries (`sql/queries/home.sql`) return album ids only and are joined back to `GetAllAlbumsWithDetails` in Go, so the album projection has one definition. A shelf with nothing behind it is omitted, never rendered empty — **and so is a shelf that repeats the one above it**, which is the same rule one step further: "On repeat" was "Pick up where you left off" reordered, because a small library has one signal and answers several questions with the same albums. Two guards make that safe, and both were arrived at by breaking the existing tests: only shelves of three or more albums are judged (two rows of one overlap by 100% whenever they agree at all), and only when the shelf is **not showing the whole library** — a repeat is a fault only if a different row was possible. Measured against a fixed shelf size instead, an 11-album library kept three identical shelves while a 13-album one lost them. **The app lands here**, from `index.ts` after the stores are wired. `index.html` still renders the track list eagerly and it is still what paints first — it is the cached `tracks` view, so the navigation is a class toggle plus one chunk rather than a second render of the shell. `app-sidebar`'s default `activeView` is `home` to match, because the sidebar does not hear a `navigate` it did not send. - `profiling` — pprof server on `:6060`, compiled out in non-dev builds via build tags (`internal/dev/`). **Explore catalog** (`backend/explore/`): the searchable MusicBrainz/ ListenBrainz catalog in `explore_index`. Deriving it from the MetaBrainz dumps means streaming ~89 GB from a server that caps a client near 2 MB/s — half a day, for a catalog identical for every user. So that work happens **once, centrally**, and users download the result: - `cmd/indexbuild` builds the catalog from the dumps; `cmd/indexexport` cuts it down to a shippable core and stamps its provenance. `.gitea/workflows/index-artifact.yml` runs both and publishes the compressed artifact under a fixed `latest` version. - The app fetches and merges that artifact (`artifactfetch.go`, `artifactimport.go`) — about a minute, versus a day. - Everything the app does **not** need is behind the `indexbuild` build tag (`dumpimport.go`, `dumpcounts.go`, `dumpcatalog.go`, `dumpproject.go`, `dumpparallel.go`, `indexpatch.go`) so it is not linked into the binary. `dumpbuild_stub.go` is the app-side entry point; `dumpshared.go` holds what both sides use. - The app keeps popularity current with the daily incremental dumps (`dumpincremental.go`), and resolves artists outside the artifact's coverage lazily on first view. **Frontend** (`frontend/`): Lit 3.2 web components + Web Awesome UI library + HTMX. State management via singleton reactive stores in `src/store/`. Wails bindings auto-generated in `frontend/wailsjs/` — don't edit by hand. **A view is a chunk, and three components are not.** `index.ts` holds a loader table (`VIEW_LOADERS`, `DETAIL_LOADERS`) and `await`s a view's module before creating its element — `document.createElement` on an undefined tag yields an inert `HTMLElement` rather than throwing, so a missing entry is a blank page, not an error. Navigations are numbered and anything after the `await` re-checks it is still the newest, or a slow chunk lands on top of a faster navigation. Every chunk is then warmed on idle, so the split is paid once at startup rather than on every first visit. **`notification-host`, `inline-notice` and `confirm-dialog` stay eager on purpose**: a failure surface that has to fetch a chunk before it can speak is not a failure surface, and the moment it is most needed is the likeliest moment loading one fails. `first-run-wizard` and the startup chrome are eager for the ordinary reason — they are the first paint. **A primary view is cached, not unmounted.** `index.ts` keeps every primary view in the DOM and toggles a `.view-hidden` class, because that is what preserves `scrollTop` across navigation — so `disconnectedCallback` never fires for one, and anything registered there runs for the life of the session from pages it is not on. The missing half is `utils/view-lifecycle.ts`: navigation calls `viewDeactivated()` on the outgoing view and `viewActivated()` on the incoming one, and a view registers its document listeners, timers and backend subscriptions through `listenWhileActive` / `intervalWhileActive` / `whileActive`, which are torn down on the way out. An off-screen view also does not render. A shared reactive controller gets the same treatment via `registerViewAware`. **The player's position comes from the player.** `seek-bar` renders `PlaybackPositionChanged` (payload `player.PositionInfo`), emitted at 1 Hz while playing and immediately on load, play, pause, seek and natural finish. Its local `setInterval` is interpolation *between* reports only, stopped and restarted by every one of them — it used to be the clock, and counted itself 30 s away from the backend across four keyboard seeks. A report carries `trackChangeId` (the store is a singleton, so a bar mounting later must not adopt a report about the previous track) and a `seq` (the same second reported twice still has to reset the interpolation). What the player cannot do, it says: `PlaybackFailed` is emitted from both the load and the play path, auto-advance **skips** the failed track (bounded by the queue length, so a disconnected drive stops after one pass), and the bottom bar shows one coalescing line — "Skipped 12 tracks that could not be played." That line is the Inline level of the app's one notification surface, below. **Failure has one voice, and the caller picks how loud.** `store/notification-store.ts` is the only notification surface; before it, 84 `catch` blocks ended at `console.error` and two components had grown private toasts. Four levels, chosen by the call site from one rule — *a failure is only worth interrupting for if the user can do something about it that they are not already doing*: - **Blocking** (`wa-dialog`, must be acknowledged) for data at risk: a folder left holding a mix of old and new tags. Two callers are anticipated; a third should be argued for. - **Persistent** (stays, with an action) for something the user asked for that did not happen and retrying is meaningful. - **Transient** (a toast) for a small action whose state visibly reverted anyway — a favourite that came back. - **Inline**, rendered by `` in the panel that failed, never as a toast. Three things about it are load-bearing. **Coalescing lives in the store**, keyed by `(level, region, key)` within a window, so 200 unplayable files are one message with a count and no future caller has to remember that. **An inline notification carries a region**, because "inline" says *not global*, not *where*. And **the bottom band belongs to the player** — the app-level stack sits under the header, since the player's own floating notice grows upward by however many lines it needs and a bottom-anchored stack collides with it on a small window. What reaches a person is a sentence: `utils/describe-error.ts` maps the causes a user can act on (offline, timeout, not found, permission, database busy) to copy, `explainError` repeats a backend message when it is one of *our* sentinels rather than a Go wrapping chain, and the raw text stays in `console.error`. The one documented exception is a download client's connection test, whose verbatim error is the user's debugging tool. Destructive actions ask once, through `confirmAction()` (`components/confirm-dialog/`), which is a `wa-dialog` and so brings the focus trap and Escape the hand-rolled overlays do not have. **Every dialog in the app is a `wa-dialog`, and there is no sixth pattern.** The four hand-rolled autotag overlays and the remove-library confirmation had no `role`, no `aria-modal`, no focus trap and no focus restore — including the two gating an irreversible on-disk metadata rewrite. The split is by *shape*, not by owner: a dialog that only asks a question is a `confirmAction()` call (title, message, impact, confirm/cancel), and a dialog carrying **input** is a `` in the host's own template. Both remaining autotag dialogs render unconditionally with `?open` deciding which is up — mounting one on demand puts the element and its `showModal()` in the same update. `autotag-view`'s last document keydown listener died with them; it existed only because its dialogs could not close themselves. **None of them had an accessible name, and one helper gives all of them one.** Every call site passes `label`; Web Awesome renders it into an `

` in the same shadow root as the native `` and never points `aria-labelledby` at it — so for eleven dialogs `getByRole('dialog', {name})` matched nothing and a screen reader announced an unnamed dialog. `utils/name-dialog.ts` sets that IDREF (and falls back to `aria-label` under `without-header`, which renders no heading to point at), called from each host's `updated()`. Three things about it are load-bearing. It **reaches into another library's shadow root**, which is open but is not API — acceptable here only because the failure is bounded: if Web Awesome moves the structure the query misses, nothing is written, and the dialog is as unnamed as it was. It uses **`aria-labelledby`, not `aria-label`**, because three call sites compute their label at render time and an IDREF to the heading Web Awesome re-renders stays correct with nothing resyncing it. And it **waits for the dialog's own first update**, not its host's: `wa-dialog` is a Lit element whose shadow root is populated in *its* update, so a query at the host's `firstUpdated` finds an empty root and names nothing — the same lifecycle trap that hid `wa-dropdown-item`'s role from the menu keyboard model. Two awkwardnesses remain, and they are about *locating* one rather than naming it. The host is `display: contents`, so the element carrying the testid always reports hidden — what is visible is the `` inside it — and what holds the slotted content is the *host's* shadow root, not the dialog's subtree. A third is worth knowing before checking any of this: the Playwright **a11y snapshot never prints a dialog's name**, named or not, so it cannot tell you whether this works. `getByRole` can, and CDP's `Accessibility.getFullAXTree` gives the browser's own answer. **A disclosure is a button, and it says what it controls.** `config-section`'s header was a bare `
` with no `tabindex`, no `role` and no `aria-expanded`, and every section defaults to collapsed — so every setting in the app sat behind a control that could not be tabbed to (the audit's last Critical). It is a `