feat(harness): agent-drivable dev harness and CI that gates
Build & publish Arch package / arch-package (push) Successful in 2m8s
CI / check (push) Failing after 1m56s
CI / e2e (push) Skipped
Search index maintenance / maintain-index (push) Successful in 13s

A coding agent could develop this repo's Go packages and could not
develop the application: every path to running YellowJacket ended in a
blocking GTK window, so 265 bound methods, 46 events, 33 component
directories and 13 stores had exactly one form of verification
available — `tsc --noEmit`.

The unlock is that `wails dev`'s dev server on :34115 serves the real
frontend with the real generated bindings against the same Go backend a
desktop window attaches to, so a plain Chromium under Xvfb gets a fully
functional app. Four test tiers now exist, cheapest first:

- `make ui-test` — 313 Vitest tests in a real browser in ~2 s, no app,
  no backend, no display. Works because `frontend/wailsjs/` is a pure
  passthrough to `window.go`/`window.runtime`, so faking just those two
  globals runs the real bindings and the real store code.
- `make test` — services in-process, asserting on the payload the
  frontend would receive, via a new `events.Emit` wrapper.
- `make dev-headless` + `playwright-cli` — the real app, driven
  interactively, with an event bridge on `window.__yjEvents` and a
  dev-only control surface at `/__test/`.
- `make e2e` — 19 of those flows frozen as Playwright specs.

`events.Emit(ctx, …)` replaces all 35 direct `runtime.EventsEmit` call
sites: wails' `getEvents` `log.Fatalf`s on any context without its
runtime, so those paths could not run under test and a background
worker could take the app down. Four packages had each hand-rolled the
same guard; nine more guarded on `ctx != nil`, which does not help.
`TestNoDirectRuntimeEmits` fails the build on a new one.

Fixtures are generated, not committed (`make testdata`), and seeds are
built by *running the app* — never by hand-writing config and DB rows,
which would be a second description of a valid YJ_HOME.

`.gitea/workflows/ci.yml` is the first workflow here that tests
anything; the other three only package, so `gitea_ci` reported only
packaging jobs and misled anyone asking whether a push was healthy.
Both jobs were prototyped to green in a bare ubuntu:24.04 container
before the YAML was written, which immediately caught `make lint`
linting three configurations that nothing builds: all three passes
omitted `webkit2_41`, so wails resolved webkit2gtk-4.0 — which Arch
still ships and Ubuntu 24.04 dropped.

Operational instructions live in `.pi/skills/yellowjacket-dev/`,
measured discoveries in `.planning/NOTES.md`, and architecture in
`CLAUDE.md` — split by tense, not by topic, because a topical split
gives every new fact two plausible homes. `make skill-check` fails a
commit if the skill cites a make target that does not exist.
This commit is contained in:
2026-08-10 23:20:42 -04:00
parent 65333857e2
commit 5ca6cad45a
117 changed files with 14585 additions and 262 deletions
@@ -0,0 +1,67 @@
# The fixture library
`test_data/music_library_test/` is **generated, not committed**:
`make testdata` (~1 s) builds 31 deterministic tracks across MP3, FLAC,
Ogg Vorbis and WAV. `make testdata-force` rebuilds unconditionally,
`make testdata-clean` deletes it. `make test` and `make sandbox-seed`
depend on it, so it is rarely run by hand.
Tests that need it fetch it through `internal/testfixtures` and skip
themselves when it has not been generated.
## Select by case, never by path
```go
m := testfixtures.Load(t)
paths := m.Case(t, testfixtures.CaseCoverDedup)
track := m.Track(t, rel)
```
Cases: `cover-dedup`, `multi-disc`, `various-artists`, `flac-album`,
`ogg-album`, `wav-tracks`, `partial-tags`, `unicode`, `duplicates`,
`edge-lengths`, `broken`.
Two invariants worth not breaking:
- **The clean library is exactly 31 tracks**, because `sandbox-seed`
verifies the scan against that count. Deliberately malformed files
live in a *sibling* root, `test_data/music_library_broken/`
(`m.BrokenPath()`), so the scanner never sees them.
- **Tags are written by `backend/tagwriter`, not by ffmpeg** (which
encodes with `-map_metadata -1`). Fixture and reader therefore cannot
drift into agreeing with each other and disagreeing with reality.
The manifest (`test_data/music_library_test.manifest.json`, outside the
scanned root) hashes the *spec* — paths, formats, durations, tags,
cover identity — not the bytes, because ffmpeg stamps encoder version
strings and identical specs produce different bytes on different builds.
## In e2e specs
- **Every fixture except one is 26 seconds.** A spec that starts
playback and then clicks pause races the track ending and fails
against a correct UI. Use `LONG_TRACK` (90 s, `edge-lengths`) exported
from `e2e/support/fixtures.ts`.
- **WAV tracks scan in untitled.** `backend/tagwriter` writes WAV tags
into a RIFF `id3 ` chunk and `dhowden/tag` has no RIFF parser, so
there is no "Field Recordings" artist in the Artists view. This is a
known open bug pinned by `TestWAVTagsAreNotReadableYet`; do not
"fix" a spec by asserting the broken behaviour elsewhere.
## Seeds
```bash
make sandbox-seed NAME=default # build (boots a fresh YJ_HOME and drives the app)
make sandbox-seeds # list
make dev-headless SEED=default # restore into a run
```
A seed is a tarred `YJ_HOME` produced by *running the app*: fresh home →
real `AddLibrary` binding → poll until the real scan reports the
manifest's track count → SIGTERM so shutdown hooks persist state → tar.
Never hand-write one. Seeding points `YJ_CORE_INDEX_URL` at a dead
address on purpose, so no seed depends on what the explore artifact
server happened to be serving.
Rebuild a seed after any schema change, or the restored database is
migrated on open in a way the seed's author never saw.
@@ -0,0 +1,85 @@
# The harness: event bridge and control surface
Two things ride on top of the headless app. Both exist only in dev
builds; neither is reachable from a shipped binary.
## The event bridge (`.playwright/init-events.js`)
Loaded as an `initScript` by `.playwright/cli.config.json` and by
`e2e/support/fixtures.ts`, so an exploratory session and a committed
spec see an identical page. It records every backend event by wrapping
`window.wails.EventsNotify` — the single choke point all 46 events pass
through, whether or not the app subscribes to them.
```js
window.__yjEvents.wait('LibraryScanComplete', { timeoutMs: 60000 })
window.__yjEvents.names() // name -> count; use this to find out
// what actually fired before asserting
window.__yjEvents.last('QueueChanged')
window.__yjEvents.since(seq)
window.__yjEvents.reset() // drop the buffer
window.__yjEvents.ready(20000) // resolves when a binding round-trips,
// which is later than DOM-ready and true
window.__yjEvents.call('queue.Queue.GetState', [], 5000)
```
- **`wait` resolves against already-buffered events as well as future
ones**, so there is no race between doing the thing and listening.
- **Install exactly one recorder.** Listeners survive across `eval`
calls; a second recorder double-counts. Call `reset()`, never
re-register.
- **`call` times out on purpose.** A binding with wrong argument types
never fires its callback. A 5 s rejection naming `.dev/app.log` beats
an infinite hang.
In specs, use the wrappers rather than `page.evaluate`:
`waitForEvent`, `resetEvents`, `eventNames`, `callBinding`, and the
`app` fixture (a page with the bridge installed and the backend
actually answering) from `e2e/support/fixtures.ts`.
## The control surface (`backend/testctl`, mounted at `/__test/`)
Gated twice: behind the `dev` build tag (with a no-op `!dev` twin) and
behind `YJ_TESTCTL=1`, which `scripts/dev-headless.sh` sets and
`make dev` does not.
| Endpoint | Use |
|---|---|
| `GET /__test/health` | is this a seeded dev build, and which library |
| `POST /__test/db/snapshot?name=X` | save the SQLite state |
| `POST /__test/db/restore?name=X` | put it back (see below) |
| `POST /__test/emit` `{name, data}` | force any backend event |
| `POST /__test/sql` `{sql, args}` | read rows, or a write count |
`TestCtl` in `e2e/support/fixtures.ts` is the typed client.
- **`emit` is the fast way to render a push-driven view** without
staging the work that would produce it — job progress, download
progress, scan progress. It calls `events.Deliver`, which *errors*
when the event reaches nobody, so a `200` means it really arrived.
- **`restore` is slow** (~40 s in the suite) because it copies every
table. Prefer snapshotting once and restoring only when a spec
genuinely mutates state.
## Traps in the config
- **The two path keys in `.playwright/cli.config.json` resolve
differently.** `initScript` is relative to the *config file's*
directory (`"init-events.js"`, not `".playwright/init-events.js"`);
`outputDir` is relative to the *shell's cwd*. Set `outputDir` to
`".playwright-cli"` and run `playwright-cli` from the repo root, or
snapshots land somewhere neither `.gitignore` nor your next `ls`
will find, and you will read a stale one from a previous session
and think a component regressed.
- **`snapshot` writes a file, it does not print the tree.** The
command prints a path under `outputDir`; read that. Only the tail
is echoed.
- **Three separate browser caches.** `playwright-cli`, `@playwright/test`
(`make e2e-setup`) and the Vitest provider (`make ui-setup`) each
download their own Chromium. One working is no guarantee for the next.
- **`getByRole('button', { name })` matches substrings.** "Play" also
matches "Add queue to playlist"; transport controls need
`exact: true`.
- **`e2e/` is its own npm package** with `"type": "module"`. Without
that, Playwright transpiles the specs to CJS and every `import.meta`
throws — reported, unhelpfully, as "No tests found".
@@ -0,0 +1,52 @@
# Changing the database schema
The reasoning — why there are two files, what the old 48-step migration
chain got wrong, and when squashing is legitimate — is in `CLAUDE.md`
under *Backend packages → database*. Read it once. This is the
checklist.
A schema change needs **two** files, not one:
1. **`backend/database/sql/schemas/*.sql`** — `CREATE TABLE ... IF NOT
EXISTS`, the literal target shape, what sqlc reads and what a fresh
install gets verbatim. Add the new column **last** in the
`CREATE TABLE`.
2. **`backend/database/sql/migrations/NNNN_description.sql`** — the
`ALTER TABLE ... ADD COLUMN` (and any index on it) that gets an
existing database to the same shape. Schema files are a no-op against
a table that already exists, so without this an upgrade never gets
the column.
Then:
```bash
make generate # sqlc + templ
go test -tags webkit2_41 ./backend/database/ # migration + column-order tests
make test
```
Rebuild any seed you rely on (`make sandbox-seed NAME=default`) and
delete your own dev `YJ_HOME` if you want to see the fresh-install path
rather than the migrated one.
## The three ways this goes wrong
- **Column order must match between the two paths.** `ADD COLUMN`
always appends, so a migrated column declared anywhere but last in
`CREATE TABLE` leaves fresh and upgraded installs disagreeing on
order — and sqlc binds `SELECT *` positionally, so one of them
silently reads the wrong field.
`TestMigrations_ColumnOrderMatchesFreshInstall` is the regression test.
- **Do not put an index on a migrated column in `sql/schemas/`.**
Schema files run *before* migrations, against a database that may not
have the column yet, and the predicate fails. Declare the index in the
migration, after the `ALTER TABLE`.
- **Do not add a third description of the schema anywhere.** A
migration's `ADD COLUMN` failing with "duplicate column name" against
an already-current database is expected and tolerated, not an error to
route around.
New queries go in `backend/database/sql/queries/`; generated Go lands in
`backend/database/sql/sqlcgen/`, which is never edited by hand. Tests
use `database.NewTestDB(t)`, built by the same `applySchema` production
uses, so the two cannot diverge.
@@ -0,0 +1,74 @@
# The component and store tier (`make ui-test`)
313 tests in a real Chromium in ~2 s with no Wails, no backend, no
seeded library and no virtual display. This is the cheapest coverage
available and where the bulk of UI regression belongs.
```bash
make ui-setup # once: the Vitest provider's own Chromium
make ui-test # behaviour only
make ui-watch
make ui-visual # + toMatchScreenshot baselines (YJ_VISUAL=1)
make ui-visual-update # re-record them
make ui-test UI_ARGS='store/queue' # filter
```
## How it works
`frontend/wailsjs/` is a pure passthrough — every binding is
`window.go[svc][Type][Method](args)`, every runtime call is
`window.runtime.X(...)`. So `frontend/test/support/wails-fake.ts`
replaces **those two globals and nothing else**, and the tests then
exercise the *real* generated bindings and the *real* store code. No
module mocking, and no second description of the Wails layer.
```ts
emit(Events.QueueChanged, payload); // push a backend event
stub('queue.Queue.GetState', state); // a value, or a function of the args
stubFailure('queue.Queue.SetQueue'); // reject, as a Go error does
calls('queue.Queue.SetQueue'); // what the frontend called back with
lastArgs('queue.Queue.SetQueue');
const el = await fixture('now-playing'); // mount; shadow()/text() query it
```
The dispatcher mirrors wails' own `desktop/events.js`, including
`maxCallbacks` expiry and the fact that a frontend `EventsEmit`
notifies local listeners *before* Go.
## Four things that will cost you time
- **Store singletons are constructed at module import**, before any test
can stub. `test/setup.ts` therefore carries import-time defaults for
the stores that read config in their constructor. Without one, a store
caches `undefined` where Go would have sent `[]`, and components crash
on `.length` — which reads exactly like a component bug and is not.
Adding a store that reads config on construction means adding its
default there.
- **`vitest.config.mts`, not `.ts`** — it `mergeConfig`s the repo's
`vite.config.mts` to reuse the `@go`/`@store`/`@components` aliases,
and a `.ts` sibling cannot import it.
- **Screenshots need the theme.** The setup file imports
`@store/theme-store` for its side effect (it applies the `--yj-*`
ramp to `:root`); without it a component renders white-on-white and
the baseline is blank.
- **`@lit-labs/virtualizer` never produces two identical frames**, so
`toMatchScreenshot` on `<queue-panel>` fails with "could not capture a
stable screenshot" rather than a diff. Assert on its rows instead.
Visual baselines are font-hinting and compositing sensitive, which is
why they are opt-in: they only mean anything on the machine that
recorded them.
## Bindings
`frontend/wailsjs/` is generated by `wails`, **not** by `go generate`,
so the pre-commit codegen check does not cover it — a renamed Go bound
method first shows up at runtime, as a call that never settles.
```bash
make bindings-check # ~1.5 s, also a pre-commit hook
make bindings # regenerate for real
```
The generator rewrites `wailsjs/runtime/*` as mode 755 every run; that
is churn, not drift, and the check ignores it.