A coding agent could develop this repo's Go packages and could not develop the application: every path to running YellowJacket ended in a blocking GTK window, so 265 bound methods, 46 events, 33 component directories and 13 stores had exactly one form of verification available — `tsc --noEmit`. The unlock is that `wails dev`'s dev server on :34115 serves the real frontend with the real generated bindings against the same Go backend a desktop window attaches to, so a plain Chromium under Xvfb gets a fully functional app. Four test tiers now exist, cheapest first: - `make ui-test` — 313 Vitest tests in a real browser in ~2 s, no app, no backend, no display. Works because `frontend/wailsjs/` is a pure passthrough to `window.go`/`window.runtime`, so faking just those two globals runs the real bindings and the real store code. - `make test` — services in-process, asserting on the payload the frontend would receive, via a new `events.Emit` wrapper. - `make dev-headless` + `playwright-cli` — the real app, driven interactively, with an event bridge on `window.__yjEvents` and a dev-only control surface at `/__test/`. - `make e2e` — 19 of those flows frozen as Playwright specs. `events.Emit(ctx, …)` replaces all 35 direct `runtime.EventsEmit` call sites: wails' `getEvents` `log.Fatalf`s on any context without its runtime, so those paths could not run under test and a background worker could take the app down. Four packages had each hand-rolled the same guard; nine more guarded on `ctx != nil`, which does not help. `TestNoDirectRuntimeEmits` fails the build on a new one. Fixtures are generated, not committed (`make testdata`), and seeds are built by *running the app* — never by hand-writing config and DB rows, which would be a second description of a valid YJ_HOME. `.gitea/workflows/ci.yml` is the first workflow here that tests anything; the other three only package, so `gitea_ci` reported only packaging jobs and misled anyone asking whether a push was healthy. Both jobs were prototyped to green in a bare ubuntu:24.04 container before the YAML was written, which immediately caught `make lint` linting three configurations that nothing builds: all three passes omitted `webkit2_41`, so wails resolved webkit2gtk-4.0 — which Arch still ships and Ubuntu 24.04 dropped. Operational instructions live in `.pi/skills/yellowjacket-dev/`, measured discoveries in `.planning/NOTES.md`, and architecture in `CLAUDE.md` — split by tense, not by topic, because a topical split gives every new fact two plausible homes. `make skill-check` fails a commit if the skill cites a make target that does not exist.
2.4 KiB
description, argument-hint
| description | argument-hint |
|---|---|
| Promote a hand-driven playwright-cli session into a committed spec in e2e/ | [name of the flow] |
Promote the flow I just drove by hand into a committed Playwright spec. Flow: ${@:-infer it from the playwright-cli commands in this session}
This is a transcription with fixed substitutions, not a fresh test. Work from what actually happened in this session, not from what the UI looks like it should do.
1. Recover the flow. List the playwright-cli calls made this
session, in order, and the assertion each one was really checking. Then,
before writing anything, ask the running app what fired:
playwright-cli -s=yj eval "() => window.__yjEvents.names()"
Await the events that are actually in that list. Do not guess event
names from backend/events/.
2. Substitute, one for one.
click e15→ a role ordata-testidselector. Snapshot refs are per-snapshot and meaningless in a spec. If the only stable selector would be structural, add adata-testidto the Lit component and re-runmake ui-test.- any sleep, or "it looked settled" →
waitForEvent(app, 'X'). window.go.…→callBinding(app, path, args), which times out.- a short fixture track →
LONG_TRACK, if the flow needs playback to still be running on the next line. Every other fixture is 2–6 s. getByRole('button', { name })→ addexact: true.
3. Place it. e2e/specs/<area>.spec.ts, importing test, expect
and the helpers from ../support/fixtures.js — never @playwright/test
directly. Match the surrounding specs' comment style: say what the test
is protecting against, not what the lines do.
4. Prove it is a spec and not a recording. Three runs, in order:
make e2e E2E_ARGS='--grep "<name>"' # it passes
make e2e E2E_ARGS='--grep "<name>"' # again — catches dependence on
# state the first run left
then once more after restoring the database through /__test/, which
catches dependence on state my hand-driving left behind — the single
most likely way a promoted spec passes here and fails in CI. Leave the
database as you found it: snapshot/restore, or reset in beforeEach.
5. Then the whole suite: make e2e. If the promotion turned up a
new trap, append it to .planning/NOTES.md; if it turned up a bug,
tell me rather than asserting the broken behaviour.