Compare commits

...
Author SHA1 Message Date
logan e897364a73 fix(loop): worktree provisioning, fresh fetch, corruption halt
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m18s
CI / e2e (pull_request) Successful in 11m3s
Three P1 adoption-wave findings, all of which break an unattended run.
The worktree needs make build-frontend and make testdata once before its
first push — the pre-push go-test hook embeds frontend/dist and the
fixture library and refused the push without them. The merge leg now
says to fetch origin/main in the same breath as the refresh merge: a
cached origin/main merged against the wrong base, CI went green on it,
and the merge came back 405 "behind base" — a full cycle wasted. And a
killed fetch leaves 0-byte objects in the shared store whose signature
(error: object file … is empty / unpack-objects failed) must halt the
loop for a human instead of churning, with the repair recipe written
down.

Closes #240
2026-09-03 13:56:58 -04:00
logan 5fa68fcf43 Merge pull request 'ci(skill-check): scan the docs a contributor reads' (#229) from docs/220-skill-check-scope into main
CI / check (push) Successful in 3m6s
CI / e2e (push) Successful in 10m59s
default
2026-09-03 17:35:40 +00:00
logan b1368bbc7e Merge remote-tracking branch 'origin/main' into docs/220-skill-check-scope
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m10s
CI / e2e (pull_request) Successful in 11m6s
2026-09-03 13:20:23 -04:00
logan 9432f68c8b Merge remote-tracking branch 'origin/main' into docs/220-skill-check-scope
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m11s
CI / e2e (pull_request) Successful in 10m44s
2026-09-03 13:04:17 -04:00
logan 4c921ed1ba Merge pull request 'fix(config): put the old value back when a setter is rejected' (#233) from fix/231-setter-rollback into main
CI / check (push) Successful in 3m52s
CI / e2e (push) Successful in 11m13s
default
2026-09-03 16:48:05 +00:00
logan f8800ca1f8 Merge remote-tracking branch 'origin/main' into fix/231-setter-rollback
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m12s
CI / e2e (pull_request) Successful in 13m24s
2026-09-03 12:29:50 -04:00
logan dddc8aaf55 Merge pull request 'fix(settings): stop offering a column the backend rejects' (#232) from fix/197-duplicate-column-label into main
CI / check (push) Successful in 3m7s
CI / e2e (push) Successful in 11m3s
default
2026-09-03 16:01:50 +00:00
logan f6e9df2f68 Merge remote-tracking branch 'origin/main' into fix/197-duplicate-column-label
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m7s
CI / e2e (pull_request) Successful in 10m54s
2026-09-03 11:44:45 -04:00
logan 47f65dad89 Merge pull request 'fix(shell): dismiss the wizard when a library exists' (#234) from fix/175-wizard-follows-the-library into main
CI / check (push) Successful in 3m17s
CI / e2e (push) Successful in 11m51s
default
2026-09-03 15:28:32 +00:00
logan f5dae71050 Merge remote-tracking branch 'origin/main' into fix/175-wizard-follows-the-library
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 4m31s
CI / e2e (pull_request) Successful in 12m3s
2026-09-03 11:09:01 -04:00
logan b0bda625e0 Merge pull request 'test(download): write the yt-dlp stub under ForkLock' (#235) from fix/146-stub-etxtbsy into main
CI / check (push) Successful in 3m12s
CI / e2e (push) Successful in 11m5s
default
2026-09-03 14:54:13 +00:00
logan 19ba5f0394 Merge remote-tracking branch 'origin/main' into fix/146-stub-etxtbsy
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m59s
CI / e2e (pull_request) Successful in 11m1s
2026-09-03 10:38:47 -04:00
logan 439a6cd77b Merge pull request 'docs(skill): the WAV fixtures scan tagged, and have since #104' (#230) from docs/225-fixtures-wav-tags into main
CI / check (push) Successful in 3m20s
CI / e2e (push) Successful in 11m7s
default
2026-09-03 14:22:59 +00:00
logan a5515d1d9f Merge remote-tracking branch 'origin/main' into docs/225-fixtures-wav-tags
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m11s
CI / e2e (pull_request) Successful in 10m51s
2026-09-03 09:55:30 -04:00
logan 7b90633456 Merge pull request 'fix(loop): refresh branches before merge, watch post-merge main CI' (#239) from feat/238-merge-leg-refresh-watch into main
CI / check (push) Successful in 3m10s
CI / e2e (push) Successful in 11m7s
default
2026-09-03 13:53:45 +00:00
logan 7838f45ed4 fix(loop): refresh branches before merge and watch post-merge main CI
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m16s
CI / e2e (pull_request) Successful in 11m30s
Adopting the six v0-era PRs surfaced two conflict-shaped cases the merge
leg handled only by luck. Behind-main branches are refused outright by
the repo's block_on_outdated_branch protection, so the leg now refreshes
every branch against origin/main before merging — which is also where a
textual conflict should surface, as diff text the loop resolves only
where it authored the hunks, otherwise abandoning the PR to a human
with a comment. And the one guard no mergeability check provides is the
push run on main after the merge: three PRs touching the same file can
merge cleanly and contradict each other, so a red main now halts the
loop instead of the tick reporting merged and moving on.

Closes #238
2026-09-03 09:38:05 -04:00
logan 2453d717cf Merge pull request 'feat(loop): autonomous backlog loop — tracker to merged main, scheduled' (#237) from feat/236-autonomous-backlog-loop into main
CI / check (push) Successful in 3m6s
CI / e2e (push) Successful in 11m37s
default
2026-09-03 03:54:15 +00:00
logan e772f51982 feat(loop): add the autonomous backlog loop configuration
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m23s
CI / e2e (pull_request) Successful in 10m53s
The scheduled backlog runs already claimed, fixed, verified and opened
PRs one issue at a time (~50 runs), but stopped at "PR open, CI green" —
every merge and every stale branch was human work, and the pipeline
shape existed only as one prompt file. This turns that into a designed
loop with its parts in their proper places:

- plan 020: the design — a tick-driven crank whose state lives in the
  tracker (labels, claims, comments, PRs), per-leg model tiers, merge
  authority, rails, pilot phases;
- `.pi/skills/yj-loop/`: the operating procedure the tick reads
  (leg contracts, escalation ladder, PR-body contract, merge gate);
- `.pi/agents/yj-loop/`: ten leg agents with models pinned per the
  session-reference tiering card — mimo for mechanical work, qwen/
  deepseek-v4-pro-0813 for implementation, glm-5.3 for selection,
  planning and consequences review, glm-5.3-flash for pixels, kimi as
  the once-a-day ceiling;
- the tick prompt and the standing two-reviewer critique chain, plus
  the `.pi/loop/` gitignore entry and the CLAUDE.md pointer.

The switch stays where the v0's was — `.pi/schedule-prompts.json`,
gitignored, live only while the loop's pi session is open. No code
changes. Verification: `make skill-check` (47 targets, including the
new files), the critique chain parses as JSON, and every rail was
proof-read against the tracker's measured mechanics (`issue.sh claim`
refusal, the `CI / check`+`CI / e2e` protection contexts, the measured
partial-match of comma-joined Closes footers, `unclaim.yml`).

Closes #236
2026-09-02 23:32:04 -04:00
logan 49445ded77 test(download): write the yt-dlp stub under ForkLock
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m48s
CI / e2e (pull_request) Successful in 10m21s
The kernel refuses to exec a file that is open for writing anywhere in
the process, and these tests are parallel: a sibling's fork duplicates
stubYtDlp's write descriptor in the moment it is open and carries it
past our close, so the exec a moment later fails with ETXTBSY. That is
the flake seen once locally and once in CI, both times on a tree with
no Go in its diff.

Closing sooner is not available -- os.WriteFile has already closed the
file before anything execs it -- and O_CLOEXEC does not help, because
the window is between another goroutine's fork and its own exec.
syscall.ForkLock is the lock forkExec takes across that fork, so
holding it over the write means no child can exist while the
descriptor does.

Measured on the helper itself under 12 concurrent writers: 176-189 of
2400 execs refused before, 0 of 2400 after, three runs each.

Closes #146
2026-08-30 07:40:32 -04:00
logan b9e60bdb0a fix(shell): dismiss the wizard when a library exists
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m58s
CI / e2e (pull_request) Successful in 10m25s
The first-run wizard dismissed itself as a step in its own flow, so a
library appearing by any other route left a full-screen modal up over
an app that was already set up — intercepting every pointer event,
which is also what keeps Settings out of reach while it is there.

AddLibrary emits LibraryAdded whoever calls it, so one subscription
makes the dismissal follow the state the wizard exists to wait for
rather than the button being pressed. It is registered before the
initial read, or a library arriving while that call is in flight is
answered with a stale empty list; both routes out now end in one
dismiss(), which sets `finished` before asking the dialog to close
because preventClose cancels the hide otherwise.

Closes #175
2026-08-30 06:40:32 -04:00
logan bcf3856b6f fix(config): put the old value back when a setter is rejected
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m46s
CI / e2e (pull_request) Successful in 10m7s
Config.Save() validates the whole config, so a setter that assigned
before validating did not merely fail its own call: the rejected value
stayed in memory and failed every later save, of every unrelated
setting, silently and for the rest of the session. Nothing reached
disk, so a restart cleared it — which is what made the fault invisible
and unreportable.

The defect is precisely "assignment precedes a validation that can
reject that argument", and that predicate enumerates seven setters
rather than the whole file. Each snapshots the field and restores it on
the error path.

The remaining setters were read rather than assumed and are unchanged:
shortcuts.Config.Validate returns nil unconditionally, the bools and
SetFavoritesPlaylistID pass through no validation that inspects them,
Config.Validate does not validate Downloads at all, and SetViewVisible
refuses an unknown, non-hideable or launch-page view before assigning.
SetLibraryDirectory was already correct and is the precedent the new
comment points at: it validates a candidate before assigning, so there
is nothing to undo.

The rationale sits above the setter section rather than on Save(),
which is bound — a doc comment there renders into frontend/bindings
for an audience with no use for it.

Closes #231
2026-08-30 05:40:36 -04:00
logan d225f922fb fix(settings): stop offering a column the backend rejects
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m50s
CI / e2e (pull_request) Successful in 10m11s
Settings -> Track List Columns listed two rows both called "Track
Name", one of which could not be ticked, and a screen reader heard
"Show the Track Name column" twice with nothing to tell them apart.

They are `titleArtist` and `trackName`, and the issue left open which
way to fix it: a second label so the sort dropdown and the configurator
can say different words, or drop the row because the user cannot select
it. The code settles that. `titleArtist` is not in
`tracklist.AllColumnIDs`, so `Config.Validate()` returns `unknown
track-list column ID: "titleArtist"` -- ticking the row sends a column
set the backend rejects, `config-page` swallows the rejection into a
`console.error`, and the tick reverts. A second label would have named
a control that cannot work.

So a definition says whether it is a *choice*, and the configurator
reads `CONFIGURABLE_COLUMN_IDS` rather than `Object.keys(COLUMN_DEFS)`.

The new test reads Go's own list out of `backend/tracklist/config.go`
rather than writing it down a third time, since a third copy is the
fault one step earlier. Verified by planting: with the filter removed
all four assertions fail with the defect's own numbers.

Not fixed here, and filed as #231: `SetTrackListColumns` assigns before
it validates, so a rejected list stays in memory and `Save()` validates
the whole config -- one tick and no setting saves for the rest of the
session. That is reachable from any invalid input, not from this row.

Closes #197
2026-08-30 04:41:54 -04:00
logan dfb338fc37 docs(skill): the WAV fixtures scan tagged, and have since #104
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m45s
CI / e2e (pull_request) Successful in 10m12s
`fixtures.md` told an agent the WAV fixtures scan in untitled, that
there is no "Field Recordings" artist in the Artists view, and that
this is a known open bug "pinned by TestWAVTagsAreNotReadableYet" — a
test #104 deleted, because it existed to assert the reader did not work
and failed the moment it did.

That last clause is why this is worth a diff rather than being left to
rot: the paragraph is an instruction, and it instructs the next reader
that a spec asserting the *working* behaviour is the mistake. It is the
same #104 staleness #217 removed from `queue-selection.spec.ts`, one
file over, still telling agents to put it back.

Measured against a running app rather than corrected from the issue
text — and the seed had to be rebuilt first, since the one on disk
predated #104 and would have replayed a pre-#104 scan and confirmed the
stale paragraph. On a fresh `make sandbox-seed NAME=default`, both WAVs
carry a title, an artist credit and an album: "Field Recordings" is an
ordinary artist with 2 tracks and "Test Tones" has a cover row. The
only two tracks with no album at all are `unsorted/no-tags-at-all.mp3`
and `unsorted/title-only.mp3`.

The replacement also says that prose written before #104 disagrees,
because it does, and saying nothing is how the next reader reintroduces
the claim from a source this change deliberately does not touch.

Deliberately carries no `Closes` footer. #225 covers two halves, and
the second — the same staleness in two *dated* `.planning/NOTES.md`
entries — is left alone: whether measured history gets a correcting
clause is a judgement about what that file is for, which the issue
raises on purpose and this change must not settle by auto-closing it.
2026-08-30 03:36:54 -04:00
logan 26251badda ci(skill-check): scan the docs a contributor reads
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m0s
CI / e2e (pull_request) Successful in 10m18s
The check asserts that every make target named in a doc exists, and its
scanned set was .pi/ plus CLAUDE.md.  Since #50, CONTRIBUTING.md is the
document a *human* goes to for a build command, and it names 21 targets
that nothing verified; README.md names none today and is in for the same
reason.  The script's own header sentence is the argument — a renamed
target sends a person off the same cliff it sends an agent off.

The file list is now one `docs` variable used twice, because the failure
message carried a second copy of it and a second list is a second thing
to forget.  The `[ -d .pi ]` guard went with it: gating the whole run on
.pi/ would make the human-facing half conditional on the agent-facing
one, and an empty list is the same "nothing to scan" exit without the
coupling.

The lefthook glob is that scanned set now rather than
{Makefile,.pi/**/*.md} — #220's smaller half, and it did not fire on
CLAUDE.md either, which the script had read for months.

Verified by planting a bad target rather than by reading the diff: both
matched forms in each of the four scanned surfaces, each naming the
right file; the same two plants pass on the pre-change script; unfenced
prose still does not match; and the hook fires on a staged
CONTRIBUTING.md under the new glob where the old one skipped it.  The
count is unchanged at 47 — the set is a union — so coverage is the only
thing that moved.

Closes #220
2026-08-28 03:38:28 -04:00
logan 5e25e14994 Merge pull request 'Split the README into a user-facing landing page and CONTRIBUTING.md' (#221) from docs/50-readme-landing-page into main
CI / check (push) Skipped
CI / e2e (push) Skipped
Build & publish the Android APK / apk (push) Successful in 1m50s
Build & publish Arch package / arch-package (push) Successful in 2m44s
Attach the desktop build to the release / linux (push) Successful in 1m17s
Sync Homebrew formula / sync-formula (push) Successful in 10s
2026-08-26 16:03:22 +00:00
logan 94ccea185c Merge pull request 'fix(ui): the phone's nav sheet says when it scrolls' (#222) from fix/210-nav-sheet-scroll-affordance into main
CI / check (push) Canceled after 0s
CI / e2e (push) Canceled after 0s
2026-08-26 16:03:10 +00:00
logan d21b842d86 Merge pull request 'fix(queue): name the queue header's two older actions' (#223) from fix/170-queue-header-action-names into main
CI / check (push) Canceled after 0s
CI / e2e (push) Canceled after 0s
2026-08-26 16:03:02 +00:00
logan f79249dfba Merge pull request 'test(e2e): name the fixture tracks that really have no album' (#226) from test/217-fixture-names-in-queue-selection into main
CI / check (push) Canceled after 0s
CI / e2e (push) Canceled after 0s
2026-08-26 16:02:53 +00:00
logan f8c8d374d1 Merge pull request 'fix(riff): grow a chunk buffer with what arrives' (#224) from fix/216-riff-parse-allocation into main
CI / check (push) Canceled after 0s
CI / e2e (push) Canceled after 0s
2026-08-26 16:02:50 +00:00
logan ec4961ae50 test(e2e): name the fixture tracks that really have no album
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m47s
CI / e2e (pull_request) Successful in 10m17s
`queueSixAndOpen` filters the queue down to tracks that have an album,
because `explore-link` renders a name it cannot route as plain text and
one test clicks that name. The filter is right and unchanged; the
comment explaining it named the wrong two files.

Since #104 read a WAV's `id3 ` chunk, the two tracks under `Field
Recordings/Test Tones` are tagged, scanned and ordinary. Asked of a
seeded app rather than of the comment, exactly two tracks in the
fixture library have no album: `unsorted/no-tags-at-all.mp3` and
`unsorted/title-only.mp3`.

The clause saying which change made the old names wrong is there so the
next reader does not restore them.

Closes #217
2026-08-26 07:42:11 -04:00
logan a113b7bd62 fix(riff): grow a chunk buffer with what arrives
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 6m23s
CI / e2e (pull_request) Successful in 10m8s
Parse sized its buffer from the chunk header, which is four bytes read
off the file, so a truncated or malformed WAV declaring a 4 GB data
chunk in a 2 kB file got 4 GB from the allocator before the read
discovered there was nothing to put in it. The error was always right;
the allocation happened first.

io.CopyN into a bytes.Buffer is what ID3Chunk beside it has done since
#104, and needs nothing new: the reader stays an io.Reader and the
buffer grows with what actually arrives.

The regression test measures rather than asserts the error, because the
error is identical on a build that allocates the gigabyte. Measured on
the pre-fix build: 1,073,750,920 bytes of TotalAlloc for a 42-byte
container whose data chunk claimed 1 GiB.

Closes #216
2026-08-26 06:35:04 -04:00
logan b5bdba2f38 fix(queue): name the queue header's two older actions
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m44s
CI / e2e (pull_request) Successful in 10m13s
Clear queue and Add queue to playlist were named by a `title` attribute
and nothing else, while the close button beside them has carried an
`aria-label` since #24. They get one too.

`title` is a name, so this is the weak-name case rather than the missing
one: it is the *last* fallback in the accname order, so any content put
inside the button later silently outranks it, and a phone has no hover
to show it. The `title`s stay — on a desktop they are also the tooltip
for an icon-only control, which is a different job.

The assertion is the part worth reading. The obvious spec — `getByRole`
by name, which is what `queue-overlay.spec.ts` already does for the
close button — is **green on the broken build**: measured against the
running pre-fix app, both buttons matched. So the second test states the
property as what it is, that the name is not the tooltip: it removes the
`title` attributes and asks again, which was 0 and 0 on main and is 1
and 1 now.

Closes #170
2026-08-26 05:39:42 -04:00
logan 20c337651f fix(ui): the phone's nav sheet says when it scrolls
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m1s
CI / e2e (pull_request) Successful in 9m56s
Since #71 the phone's "More" is a bottom sheet, and at the reference
viewport it does not fit: measured at 424x439 with the seed's eight
destinations, the body is scrollHeight 412 against clientHeight 373, so
39px is below a fold nothing announces. Where the cut lands on a row
boundary the sheet ends in a clean edge that reads as the end of the
list, which is what #207 fixed one sheet over.

The rule is that sheet's, not a second answer to the same question:
#207's two background layers move into styles/sheet-scroll.css.ts and
both sheets adopt them, with the colour left to each host as
--yj-sheet-surface. The nav sheet paints the sidebar's --yj-bg-surface
and the context sheet the menus' --yj-bg-elevated, so a shared rule that
hard-coded either would draw that seam across the other one.

The half that makes it visible is that nothing inside the sheet may
repaint the surface. These are layers on the scroller, and app-sidebar's
host carries the same grey -- in the shell its own background, in the
sheet a second opaque copy of the sheet's, over the fade. With the
fragment adopted and that rule missing, the running app measured a flat
52,58,64 to the bottom edge with 39px still below: the defect unchanged,
with every assertion about background-attachment passing. menu-surface
already meets it from the other side, where the sheet's panel is
background-color: transparent.

Closes #210
2026-08-26 04:44:02 -04:00
logan 1c08d8db90 docs: split the README into a landing page and CONTRIBUTING
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 2m58s
CI / e2e (pull_request) Successful in 10m7s
The README was two documents in one, and neither reader was served by
the other's half. It opened on a feature list, then spent its second
half on Go versions, WebKitGTK packages and `make` targets — while its
install table named a `darwin-universal.app.zip` and a
`windows-amd64.exe` that nothing has ever produced, and its header
claimed Windows and never mentioned Android, which is the one platform
with a published, self-updating channel.

So this is a correctness pass as much as a friendliness one. The README
now answers a user's questions only: what the app is, three screenshots
from the seeded fixture library so anyone can retake them, the four
formats, one install section per channel that names what is actually
published, first run, where the data lives, and pointers out. The
version-restart note is linked to the two documents that own it rather
than copied, because a copy is a second thing to keep true.

CONTRIBUTING.md takes the technical half: prerequisites, the system
libraries, the build and codegen commands, which verification tier a
change demands, the tracker workflow, the commit grammar and the style
rules. CLAUDE.md is unchanged apart from one paragraph naming the split
— it was already the deep reference both of the others point at, and
stays the only one of the three that explains why a shape is what it is.

Closes #50
2026-08-26 03:37:13 -04:00
logan 245647f12b Merge pull request 'feat(ui): warm album art ahead of the scroll' (#215) from feat/65-art-prefetch-ahead into main
CI / check (push) Successful in 2m43s
CI / e2e (push) Successful in 10m14s
2026-08-25 17:58:42 +00:00
43 changed files with 2327 additions and 160 deletions
+4 -2
View File
@@ -89,7 +89,9 @@ build/android/overlay.json
# into scripts/gitea-release.sh; the release page is the changelog.
.release-notes.md
# Agent session log: local scratch, not repo memory (that is CLAUDE.md
# and .planning/). Written by the scheduled backlog runs.
# Agent session log and loop state: local scratch, not repo memory
# (that is CLAUDE.md and .planning/). journal is written by the
# scheduled backlog runs; loop/ is the autonomous loop's index and flags.
.pi/journal.md
.pi/schedule-prompts.json
.pi/loop/
+25
View File
@@ -0,0 +1,25 @@
---
name: diffreview
package: yj-loop
description: Scope-tight review of a loop PR's diff for correctness within the plan's stated scope. The understood-diff half of the critique fan-out.
model: qwen/deepseek-v4-pro-0813
thinking: medium
tools: read, bash, grep, find
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
---
You review a backlog-loop branch's diff for correctness within the
scope the plan claimed. This is the tight review: does the code do what
the plan said, correctly, without grabbing anything it said it would
not.
Read the issue, the plan comment, and the diff itself. Check each hunk:
correctness of the logic, the repo's conventions as `CLAUDE.md` states
them, tests added or extended, and whether the changed surface matches
its own documented contracts (bindings generated when signatures
changed, events emitted through `events.Emit`, lint grammar). Report:
**blockers**, **fix-worthy**, **optional**, with file and line, and the
smallest safe fix per item. Do not modify files. Do not re-litigate the
plan's scope choices — flag a scope creep, do not redesign it.
+29
View File
@@ -0,0 +1,29 @@
---
name: escalate
package: yj-loop
description: The loop's ceiling — re-runs a leg the two lower tiers failed, seeded with their written failure summaries. Fresh session, never parallel, once a day.
model: go/kimi-k3
thinking: max
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yellowjacket-dev
---
You are the escalation tier of the YellowJacket backlog loop. Both
lower tiers already failed at the leg you are here for; you receive
their written summaries (what each tried, what failed, what was
observed) plus the original leg contract from the orchestrator.
Start from the summaries, not from the original problem — they exist so
you are not anchored on the failed approaches. Read `CLAUDE.md` and
`.planning/NOTES.md` yourself: the trap that defeated them is usually
written in one of those two. `yellowjacket-dev` tells you how to run
the harness tiers.
You may delegate mechanical subtasks, never the leg. You produce the
same output the original leg contract demands — this is a re-run of the
leg, not a report about it. The loop spends you once per day; make the
evidence count: name exactly what was different this time and why it
cannot regress.
+33
View File
@@ -0,0 +1,33 @@
---
name: inspect
package: yj-loop
description: Mechanical gatherer for the backlog loop — dumps tracker, PR, CI and branch state verbatim into a digest. No judgement, no writes beyond the digest.
model: go/mimo-v2.5
thinking: off
tools: read, bash, grep, find
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
progress: true
---
You gather state for the YellowJacket backlog loop. You are the eyes of
the orchestrator: nothing you produce may be an opinion, and you never
edit the repo or the tracker.
Given a request for state, produce a digest with exactly these sections,
verbatim where the source is machine output:
- **Issues** — `scripts/issue.sh list | search` output as relevant.
- **Pull requests** — from the REST API, open PRs with head sha and
status.
- **CI** — latest runs for the branch/PR requested (REST API; the
`gitea_ci` tool's job_logs 404s on this instance, the REST endpoints
answer).
- **Branches** — `git ls-remote --heads origin`, grepped as asked.
- **State file** — `.pi/loop/state.json` contents, untouched.
Conventions: env `GITEA_TOKEN` is required; API base
`https://git.ljones.me/api/v1/repos/yonlu/yellowjacket`. If a source
fails, report the failure exactly — never guess its contents. Keep the
digest compact; raw output over prose.
+33
View File
@@ -0,0 +1,33 @@
---
name: plan
package: yj-loop
description: Writes the implementation plan for a claimed backlog issue, as a tracker comment. Designs on the repo's real shape, not from first principles.
model: glm/glm-5.3
thinking: high
tools: read, bash, grep, find, write
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yellowjacket-dev
---
You write the implementation plan for one claimed YellowJacket issue.
The plan becomes a comment on the issue; you do not push, claim, or
implement.
Read in order: `CLAUDE.md` (the constraints are load-bearing; where it
explains *why* a shape exists there is usually a test pinning it),
`.planning/NOTES.md` (rejected approaches are rejected forever — do not
resurrect one), `.planning/plans/active/`, `.pi/journal.md`, then the
issue and any comments on it. Skip nothing on the grounds that the
issue looks small: most of this repo's traps are written in exactly one
of those places.
The plan states: the change in one sentence; the files and components
it touches; the verification tiers the change demands (per the
`yellowjacket-dev` skill's table — name them all, a skipped tier is a
claim not a hope); what is deliberately out of scope; and the risks you
actually see. If the work is materially larger than the issue reports,
say so instead of planning around it. Keep it to a screen; the worker
reads this cold.
+28
View File
@@ -0,0 +1,28 @@
---
name: review
package: yj-loop
description: Fresh-context consequences review of a loop PR — what breaks that the diff did not say. Advisory only; findings, never edits.
model: glm/glm-5.3
thinking: medium
tools: read, bash, grep, find
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
---
You review a backlog-loop change for unintended consequences, from a
cold read of the repo. Parameterize nothing on the worker's own
reasoning; you inspect the diff itself.
Read: the issue, its plan comment, `CLAUDE.md`'s load-bearing shapes,
and the branch diff against origin/main. Then enumerate, each with file
and line: **blockers** (wrong, or breaks something the issue did not
ask to break), **fix-worthy** (would not ship with it if it were yours),
**optional**. For every fix-worthy item, the smallest safe change.
Your angles: does it violate a shape `CLAUDE.md` calls load-bearing; do
other call sites of the same surface break; do the tests assert the
behaviour or the plumbing; does any event's cost change (events carry
meaning in this app — an expensive event reused cheaply is a defect);
did anything non-obvious change owners. Do not modify files. Ignore
style dust unless it hides a bug.
+27
View File
@@ -0,0 +1,27 @@
---
name: scribe
package: yj-loop
description: The loop's clerk — commit messages, PR bodies, journal and changelog-sized entries, written from supplied facts. Prose only.
model: go/mimo-v2.5
thinking: off
tools: read, bash, write, edit
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
---
You write the loop's prose. The orchestrator supplies the facts; you
shape them; you decide nothing.
Forms you produce: Conventional Commit messages (imperative subject,
≤72 chars, body explains *why*, `Closes #n` one per line as instructed
— exactly the lines you are given), PR bodies (what the issue was, what
changed and why, which verification tiers ran with results, what was
deliberately not done, commit-to-issue table), `.pi/journal.md` entries
(facts: what was done, verified, left open), and `CLAUDE.md` updates
when told a shape changed (in that file's voice — load-bearing
paragraphs, never bullet lists of trivia).
Never invent a fact: a tier result you were not given is not run. Never
rephrase a `Closes` line. Keep every form compact; this repo's prose
density is a feature.
+34
View File
@@ -0,0 +1,34 @@
---
name: select
package: yj-loop
description: Picks the single next issue the backlog loop should take. Judgment leg on the tracker state; writes nothing to the tracker itself.
model: glm/glm-5.3
thinking: medium
tools: read, bash, grep, find
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yj-loop
- yellowjacket-dev
---
You choose which one issue the YellowJacket backlog loop works next. You
are given a fresh tracker digest. You write nothing to the tracker; the
orchestrator claims.
Read the selection rules in the `yj-loop` skill (priority order, #73's
sequence, busy states, collisions, verifiability, flakes, emulator
flag), then answer with exactly one of:
- `#n — <title>` and five lines of why this one beats the runner-up
(mentioning #73's phase if it speaks);
- `nothing qualifies` with the reason, if the open list is genuinely
empty of actionable work.
Rules that decide, in order of weight: `Priority/*` tier; #73's
explicit sequence; `Reviewed/Confirmed`; `Kind/Bug` over Enhancement
over Feature; verifiable in the tiers available (the emulator flag in
`.pi/loop/state.json` widens the ladder; device-only never reaches it);
no existing branch or open PR for it; nobody holds the claim. Pick one.
Uncertainty about the tracker state is a reason to say so, not to guess.
+30
View File
@@ -0,0 +1,30 @@
---
name: validate
package: yj-loop
description: Checks that the implemented work actually answers the issue's claim, against the acceptance evidence. Claim-first validation before any review.
model: glm/glm-5.3
thinking: medium
tools: read, bash, grep, find
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yellowjacket-dev
---
You validate one issue's implemented work — the branch diff, the
worker's handoff, and the issue itself — before review and merge.
Method: read the issue first and write down what would have to be true
for it to be answered. Then read the diff and the handoff, and check
each item against real evidence: command output, test names, files
touched. Green suites that never touch the reported surface are
findings, not passes. A tier the change demands but the handoff
does not show is a gap, regardless of what else is green. Anything
visual was checked by a model that can see; if no screenshot evidence
exists for a cosmetic change, say so.
Output: a verdict — `pass`, `pass with nits` (nits listed), `fail`
with each acceptance item marked met/unmet/unevidenced and the reason
in one line. You do not edit files. You do not trust the diff's self
description; you read it.
+27
View File
@@ -0,0 +1,27 @@
---
name: visual
package: yj-loop
description: Reads screenshots of the app for the loop — the only leg allowed to judge pixels. What the image actually shows, not what the change claims.
model: glm/glm-5.3-flash
thinking: minimal
tools: read, bash
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yellowjacket-dev
---
You are the loop's eyes. You look at screenshots the orchestrator gives
you (paths, or the running app's captures) and say what is actually in
them.
Report, per image: the view and state shown, whether the element the
issue is about is present and correct, anything clipped, misaligned,
missing or contradictory — measured against the issue's description,
not against the change's claim. Where the harness provides before/after
pairs, read the difference. Be specific in pixels.
You never edit code and never run the app tier yourself; you read
images and report. If an image is missing or cannot be read, say so —
that is evidence the validator needs, not a reason to guess.
+35
View File
@@ -0,0 +1,35 @@
---
name: work
package: yj-loop
description: The loop's implementer — builds the claimed issue from its plan comment, in the loop worktree, runs the tiers the change demands, and hands off with evidence. The single writer.
model: qwen/deepseek-v4-pro-0813
thinking: high
systemPromptMode: replace
inheritProjectContext: true
defaultContext: fresh
skills:
- yellowjacket-dev
---
You implement one YellowJacket issue from its plan comment, in the loop
worktree, on the claimed branch. You are the only writer. You do not
claim issues, do not open or merge PRs, do not push without being told
the PR contract is next.
Read in order: `CLAUDE.md`, `.planning/NOTES.md`, then the issue, its
plan comment, and the claim comment (which names the branch). Implement
what the plan says and nothing else. Match surrounding style. Follow
`CLAUDE.md`'s shapes rather than reasoning from first principles.
Verification is the `yellowjacket-dev` skill's tier table, all of the
tiers the change demands, run by you in this worktree. Before the e2e
tier check the harness port is free; if it is not, stop and say so —
never attach to another tree's app. Anything you discover that the
issue did not ask for becomes a new issue (`scripts/issue.sh new`),
never a bigger diff. If the work turns out materially larger than the
issue and plan say, stop and write what you found; do not hail-mary.
Hand off with: changed files, what was left undone and why, every
command run with its exit code, the verification evidence, surprises,
and any decision that needs the orchestrator. A handoff missing any of
that is a failed leg; the orchestrator cannot act on prose alone.
+28
View File
@@ -0,0 +1,28 @@
{
"context": "fresh",
"chain": [
{
"parallel": [
{
"agent": "yj-loop.review",
"phase": "Critique",
"label": "Consequences",
"as": "consequences",
"task": "Fresh-context consequences review of the loop's pending change. Issue, plan comment and branch: {task}. Read the issue, the plan comment, CLAUDE.md's load-bearing shapes, and the branch diff against origin/main. Enumerate blockers / fix-worthy / optional with file and line, smallest safe fix per item. Do not modify project/source files; returning findings through the configured output artifact is allowed.",
"output": "critique/consequences.md",
"outputMode": "file-only"
},
{
"agent": "yj-loop.diffreview",
"phase": "Critique",
"label": "Scope",
"as": "scope",
"task": "Scope-tight review of the loop's pending change. Issue, plan comment and branch: {task}. Read the issue, the plan comment and the diff. Does the code do what the plan said, correctly, within its claimed scope? Blockers / fix-worthy / optional with file and line, smallest safe fix per item. Do not modify project/source files; returning findings through the configured output artifact is allowed.",
"output": "critique/scope.md",
"outputMode": "file-only"
}
],
"concurrency": 2
}
]
}
+26
View File
@@ -0,0 +1,26 @@
---
description: One tick of the autonomous YellowJacket backlog loop
---
You are the orchestrator of the YellowJacket backlog loop, waking for
one tick. Work in this directory. Read `.pi/skills/yj-loop/SKILL.md`
first — it is the operating procedure and it binds you. The design
questions are answered in `.planning/plans/active/020-autonomous-backlog-loop.md`;
the skill is what you run.
One tick means:
1. Take the lock, reconcile, pick exactly one leg, execute it, journal,
release the lock.
2. Delegate every deliberative leg to its `yj-loop.*` agent by name —
the model is pinned in the agent file, never an argument. You hold
only claim, shipping polls, merge, housekeep.
3. Touch only what the loop created. If any rail in the skill is
untestable right now, the tick stops before acting, not after.
4. If the scheduler fires while you are mid-answer, finish this tick
only. Two ticks never overlap; the lock is yours.
Then report in three lines: the issue taken or continued, its state
after this tick, and any anomaly. Stop. Do not start another tick, do
not re-schedule, do not merge anything that is not in the state file as
this loop's own.
@@ -42,11 +42,16 @@ strings and identical specs produce different bytes on different builds.
playback and then clicks pause races the track ending and fails
against a correct UI. Use `LONG_TRACK` (90 s, `edge-lengths`) exported
from `e2e/support/fixtures.ts`.
- **WAV tracks scan in untitled.** `backend/tagwriter` writes WAV tags
into a RIFF `id3 ` chunk and `dhowden/tag` has no RIFF parser, so
there is no "Field Recordings" artist in the Artists view. This is a
known open bug pinned by `TestWAVTagsAreNotReadableYet`; do not
"fix" a spec by asserting the broken behaviour elsewhere.
- **WAV tracks scan like every other format.** #104 added
`backend/riff`, so the scan reads the `id3 ` chunk `backend/tagwriter`
writes and both WAVs come in fully tagged: "Field Recordings" is an
ordinary artist in the Artists view, with a "Test Tones" album and a
cover. They are therefore not an example of an untitled or albumless
track — the only two tracks with no album are
`unsorted/no-tags-at-all.mp3` and `unsorted/title-only.mp3`. Prose
written before #104 says the opposite and names
`TestWAVTagsAreNotReadableYet`, a test that change deleted; that is
dated history rather than a description of the app.
## Seeds
+288
View File
@@ -0,0 +1,288 @@
---
name: yj-loop
description: Operating the autonomous backlog loop — the crank that works the YellowJacket tracker one issue at a time (tick mechanics, the state machine in Gitea, which agent and model take each leg, the escalation ladder, merge authority and the rails that stop it doing damage). Use whenever a scheduled tick fires, and when piloting or debugging the loop.
---
# The YellowJacket backlog loop
Design and arguments: `.planning/plans/active/020-autonomous-backlog-loop.md`.
This skill is the **operating procedure**; the plan is the reasoning.
`yellowjacket-dev` is the harness doctrine (tiers, seeds, traps); this
skill is the loop doctrine (who acts, on what model, with what authority).
Read the plan first, once. Then this file every tick.
## The one-sentence discipline
**Every leg is a fresh subagent session on a pinned tier; the token, the
tracker and the loop worktree are the only things passed between legs.
Never switch a model mid-session, never let two writers exist at once,
never keep state in a conversation.**
## Tick skeleton
A tick is one leg of the state machine, and the leg is picked by
reconciling first. Execute in this order:
1. **Lock.** `/tmp/yj-loop.lock` holds `pid + start-iso`. If a live
process owns it and is younger than 2 h: exit immediately, report
"tick skipped (lock held)". If the PID is dead, take the lock.
Remove it before every exit.
2. **Reconcile.** Fresh reads, never cached: open issues
(`scripts/issue.sh list`), PRs and CI via the REST API, branches via
`git ls-remote --heads origin`, `.pi/loop/state.json`. GITEA_TOKEN
refusing = the tick reports and exits; the identity rails below are
not optional.
3. **Pick the leg.** See the state machine below; the leg follows the
issue's lifecycle (claim→plan→…→merge→…→housekeep). Exactly one leg.
4. **Execute** — the leg table below says who acts and what they must
return.
5. **Journal** — one line per tick in the state file (issue, leg, result,
tick cost if leg reports it).
6. **Report** — three lines: issue taken or continued, its state now,
anomalies. Then stop. A tick that reports is a tick that can leave a
conversation behind.
## The state machine
The tracker is the truth. The state file (`.pi/loop/state.json`,
gitignored) is an index plus flags (`emulator`, `drain`); the tracker
wins every disagreement.
| Stage | Where it lives | Leg → actor |
|---|---|---|
| selected | nothing written until claim is possible | select |
| in flight | `Status/In Progress`, assignee, comment with branch+approach | claim (orchestrator, `scripts/issue.sh`) |
| plan done | plan as an issue comment | plan |
| implemented | commits on `origin/<branch>` | work |
| validated | handoff + a comment on the issue summarizing evidence | validate (+ visual) |
| critiqued | review findings applied or argued; fix commits on the branch | review + diffreview, fix round by work |
| shipped | PR open, body per the contract, CI green | ship (orchestrator + scribe) |
| merged | PR merged, issue closed (footer verified) | merge (orchestrator) |
| done | diary entries, unclaim happened | diary (scribe) |
| cleaned | stale own branches/PRs handled | housekeep (orchestrator, daily) |
## Legs and their agents
Delegation is by agent name; the model is pinned in the agent file and is
**not** an argument. Every leg prompt names: the issue, the evidence so
far (plan comment, handoffs), what the leg must produce, and its stop
rules. Never "go fix it" — the leg contract is in this file.
| Leg | Agent | Model (tier) | Produces |
|---|---|---|---|
| gather/mechanical dump | `yj-loop.inspect` | go/mimo-v2.5 (T0) | tracker/PR/CI/branch digest, verbatim |
| select next issue | `yj-loop.select` | glm/glm-5.3 (T2) | one issue + reasons, or "nothing qualifies" |
| plan | `yj-loop.plan` | glm/glm-5.3 (T2) | a plan comment on the issue |
| implement | `yj-loop.work` | qwen/deepseek-v4-pro-0813 (T1) | commits + a handoff (see contract below) |
| validate | `yj-loop.validate` | glm/glm-5.3 (T2) | pass/fail with evidence per acceptance item |
| visual evidence | `yj-loop.visual` | glm/glm-5.3-flash (T2) | what the screenshot actually shows |
| consequences review | `yj-loop.review` | glm/glm-5.3 (T2) | blockers / fix-worthy / optional findings |
| understood-diff review | `yj-loop.diffreview` | qwen/deepseek-v4-pro-0813 (T1) | same shape, scope-tight |
| escalation | `yj-loop.escalate` | go/kimi-k3 (T3) | same leg re-run, seeded with failure summary |
| prose (PR body, commit msgs, journal) | `yj-loop.scribe` | go/mimo-v2.5 (T0) | text only, from supplied facts |
Orchestrator-only legs: **claim** (`issue.sh claim --branch` — atomic,
refuses if held), **ship's PR/CI polling** (REST API below — `gitea_ci`
job_logs 404s on this Gitea; the REST endpoints are the way), **merge**
(API below), **housekeep**.
## Selection rules (`select`)
The rules from `.pi/prompts/next-issue.md` stay — priority order, #73's
sequence overriding labels where it speaks, skipping `Status/*` states
that mean busy, branch-collision check, verifiability, flakes. The
emulator flag **adds** emulator-verifiable Android issues; it never
reaches device-only ones. A "nothing qualifies" answer is a correct
tick, not a failure — report it and stop.
## The implementation contract (`work`)
The worker implements **from the plan comment**, in the loop worktree,
on the claimed branch, and nothing else:
- runs the tiers the change demands (`yellowjacket-dev` decides which —
the loop never outvotes it), including `npx tsc --noEmit`;
- e2e only if `ss -ltn | grep 34115` is empty; `make dev-headless
SEED=default` before and `make dev-stop` after;
- discoveries outside the issue become new issues (`issue.sh new`), never
bigger diffs; a materially-larger-than-implied issue stops the leg with
a comment and a label removal, not a hail-mary;
- handoff must state: changed files, what was left undone, commands run
with exit codes, verification evidence, surprises, decisions needing
approval. A handoff without that list is a failed leg.
## Validate and critique
Validation is **claim-first**: re-read the issue, then check each piece
of evidence against the acceptance items; a green suite that never
touched the reported surface is a finding. Screenshots go to `visual`,
never to a text-only tier.
Critique is the standing fan-out (`subagent` parallel: `yj-loop.review`
consequences + `yj-loop.diffreview` scope-tight, both fresh). The
orchestrator synthesizes: blockers and fix-worthy findings go back to
`work` as one bounded fix round (maximum three rounds total; then the
issue gets a `⟦loop⟧` comment stating what will not be fixed and why,
and the ship leg proceeds unless a finding is a blocker). Reviewers do
not edit files.
## Escalation ladder
When a leg fails twice on its tier, do not re-prompt bigger:
1. The failing session writes its summary: what it tried, what failed,
what it observed.
2. A **new** session on the next tier up is seeded with that summary and
the original leg contract.
3. T3 is the ceiling: fresh session, never parallel, **once per day**.
A day's escalation is spent — the issue waits until tomorrow.
Routing down is free; routing up is the budget.
## Ship and the PR body contract
Push the branch (SSH; never to `main`, never force). The PR body —
written by `scribe` from the validator's and reviewers' output — states:
what the issue was, what changed and why, **which verification tiers ran
and their results**, what was deliberately not done, the commit-to-issue
table, and `Closes #n`. `Closes` also sits one-per-line in a commit body
**inside the branch** — both, regardless of merge strategy, because the
pairing was measured.
Poll CI until `check` and `e2e` finish. On failure: read the log via
`GET /api/v1/repos/yonlu/yellowjacket/actions/runs/<run>/jobs` (per-step)
and `…/actions/jobs/<id>/logs` (full). Fix on the branch. **Two
consecutive identical failures = stop**: comment what is known on the
PR and the issue, leave both, report. Do not burn ticks on a red wall.
## Merge authority
Merge when, and only when, **all** hold:
- the PR was opened by this loop (it is in the state file's index);
- the protection contexts `CI / check` and `CI / e2e` are green on the
PR's head, read from the API, not from the PR page's badge;
- the PR reports mergeable;
- the critique leg ran and no open blocker stands;
- the branch is **not behind `origin/main`** — the protection's
`block_on_outdated_branch: true` refuses it anyway; never
`force_manually_merged` around it.
**Refresh before every merge.** In the loop worktree: `git fetch origin`
in the same breath, then `git merge origin/main` on the PR branch,
push. The fetch must be immediate — a cached `origin/main` merges
against the wrong base, CI goes green on it, and the merge comes back
405 "behind base", one whole CI cycle wasted (measured on the adoption
wave). A textual conflict
stops the leg there — as diff text, not as a failed merge click: hunks
the loop authored are resolved by the loop; anything else is left with
`⟦loop⟧` comment for a human, never forced. After any refresh push,
re-poll the PR's own required contexts on the **new head** before
merging.
**Merges happen one at a time**, each re-reading state — the previous
merge moved `main`, and the next PR's mergeability is recomputed at
its own turn.
```
curl -sS -X POST -H "Authorization: token $GITEA_TOKEN" \
-H "Content-Type: application/json" \
https://git.ljones.me/api/v1/repos/yonlu/yellowjacket/pulls/<n>/merge \
-d '{"Do":"merge","merge_message_field":"default","force_manually_merged":false}'
```
**Afterwards watch the `push` run on `main`** — the CI the merge
started. A red main after a loop merge is a **halt**: comment what is
known on the offending PR, mark the state file, stop taking new issues.
That run is the only thing between a clean textual merge of
independently-written PRs and a self-contradicting main; no
mergeability check sees it. Only a green main lets the tick proceed (to
footer verification, below).
Footer verification: `scripts/issue.sh list --state open` and check
the footer took. Close stragglers with `issue.sh close`, naming the
merge commit. `unclaim.yml` handles the label; it is not instant;
reopening does not restore it. Merging fans out to nothing (releases
are the manual `release.yml`, which the loop never runs) — the
criticism stands before the merge because nothing stands after it.
## Rails — the loop's absolute rules
1. **Touch only its own.** Issues it claimed, branches it made, PRs it
opened. `issue.sh claim` enforces the front gate; never work around a
refusal.
2. **One writer, one issue.** The loop worktree is the only dirty tree.
3. **Never merge a PR it did not open.** Any merge that violates this is
a hard stop.
4. **Human work is holy.** Human branches, PRs, assignees: leave exactly
as found. Cleanup never names them.
5. **The token is identity.** If GITEA_TOKEN misbehaves, the tick stops.
6. **New findings are new issues**, never scope creep. The tracker
vocabulary (`Kind/`, `Area/`, `Priority/`) stays intact in one
taxonomy; use `scripts/issue.sh new` with correct labels.
7. **Conventional Commits**, enforced by `scripts/commit-check.sh`; the
type list and `.releaserc.yml`'s must agree — a loop commit is a
release grammar token even after months of no manual releases.
8. **Tiers over vibes.** `yellowjacket-dev`'s tier table decides what a
change must pass; a skipped tier is stated, never silent.
9. **Two strikes on CI, three rounds of critique, one kimi a day.** The
loop's patience is finite on purpose.
10. **Every leg writes its evidence.** A leg that leaves nothing behind
is indistinguishable from a leg that did not run — which is how the
next tick re-does it.
11. **The loop may not re-schedule itself** (the scheduler refuses it
anyway — treat as an invariant, not a limitation).
12. **Drain means drain.** `drain: true` = finish in flight, take
nothing new, then stop.
## Emulator mode
Flag `emulator: true` in the state file **and** an already-booted
emulator (`adb devices` answers) opts in: `make android` (build), `make
android-install`, `make android-smoke` (crash check — the same pid
surviving is the only signal that means started), `make
android-screenshot` and `make android-eval` as evidence for `visual`.
The loop never boots or stops an emulator; that is the user's machine.
Device-only issues stay open under either setting. One-time setup the
user performs: `make android-setup` (~3.5 GB, creates the `yj-test`
AVD), then `make android-emulator` per session.
## ON / OFF / drain
- **Worktree:** `git worktree add ~/.paseo/worktrees/loop/jumpy-hound
origin/main` (from any clone; branch from origin/main in the loop
tree, never `git checkout main`). **Provision it once before the
first push:** `make build-frontend` and `make testdata` — the pre-push
`go-test` hook needs `frontend/dist` (the `//go:embed` in `main.go`)
and the fixture library, and refuses the push without them.
- **Session:** pi in that worktree, `/name loop`. Add the job via
`/schedule-prompt` (name `yj-loop`, cron
`0 0 10-18 * * 1-5`, prompt: "Read `.pi/skills/yj-loop/SKILL.md` and
run exactly one tick. Stop.") — session-bound by default.
- **OFF:** toggle the job, or close the session. **ON:** `pi --resume
loop` in the worktree, job enabled. Courses of the tick appear in
that session's transcript.
- **Tune in:** the same resume. Talk to it only between; a tick is
atomic.
## Troubleshooting
- `issue.sh: GITEA_TOKEN is not set` or a 401 — the token is the whole
identity (rails 5). Stop, do not fall back to anything.
- `gitea_ci`'s job log 404s — the REST endpoints above answer; this is
a Gitea build, not a fault.
- A spec fails that the tier doc says can fail from stale backend state
— restart the app tier before believing it (`yellowjacket-dev`).
- A tick that "did nothing" — reconcile again; the tracker usually says
which leg it really is.
- The job did not fire — the scheduler fires only while a session is
open in its directory (documented); "the loop is off" is the correct
reading, not a bug.
- `error: object file … is empty` / `unpack-objects failed` / `bad
object refs/heads/…` during a fetch or checkout — the shared object
store was corrupted (a killed fetch leaves 0-byte object files, and a
local ref can end up pointing at the dead sha1). **Halt and report**;
do not retry, the churn only deepens it. Human repair: delete the
0-byte objects, `git fetch origin --prune`, delete any ref that
still dangles (`git update-ref -d refs/heads/<b>`), re-checkout the
worktree at `origin/main`, then `git fsck --full`.
@@ -0,0 +1,244 @@
# 020 — The autonomous backlog loop
**Issue:** #236 (`Kind/Enhancement`, `Priority/Low`)
**Status:** active — phase 0, supervised pilot
**Relates:** #73 (the roadmap the loop follows), plan 005 (the harness the
loop drives). Cost and model-tier doctrine is the `pi-session-reference`
card handed to the session that designed this; the loop's copies of it
are deliberate one-paragraph summaries, not the authority.
A pi coding-agent configuration that, toggled on, works the Gitea tracker
one issue at a time — triage, claim, plan, implement, validate, critique,
PR, CI, merge, verify-close, diary — and then does it again. The tracker is
the state machine: whoever reads Gitea sees exactly where the loop is,
which is the property this document's rails exist to protect.
---
## The shape: a crank, not a resident brain
Half the design is that **nothing lives in a conversation**. Each tick is a
fresh, bounded unit of work; every transition writes evidence to Gitea
(label, comment, branch, PR) or to the loop's own state file; a tick that
dies mid-leg loses nothing, because the next tick resumes from what Gitea
says.
The other half is that **no leg trusts the one before it**. The worker
implements from the plan, not from the issue alone; the validator checks
the *claim*, not the green CI row; the merger merges only after reading the
protection contexts itself; the diary leg is what makes the next issue's
triage cheaper.
One issue in flight at a time. That is a pacing decision, not a
concurrency limit of the tooling — CI has a capacity-1 runner and the e2e
tier owns one headless port on this machine, so two writers would serialize
on infrastructure they cannot see and appear to be doing fine.
## The state machine
| Leg | Writes | Actor / model |
|---|---|---|
| reconcile | — | orchestrator + `inspect` (mimo-v2.5) |
| select | nothing on the tracker; decision logged in the tick transcript | `select` (glm-5.3) |
| claim | assignee + `Status/In Progress` + comment naming branch & approach | `scripts/issue.sh claim` |
| plan | plan as an issue comment | `plan` (glm-5.3) |
| implement | commits on the issue branch, in the loop worktree | `work` (qwen/deepseek-v4-pro-0813) |
| validate | verification evidence in the handoff | `validate` (glm-5.3), `visual` (glm-5.3-flash) for screenshots |
| critique | review findings; fix commits | `review` (glm-5.3) + `diffreview` (qwen) + fix round by `work` |
| ship | push, PR with body contract, CI read + fixes | orchestrator + `scribe` (mimo-v2.5) |
| merge | the merge; post-merge issue verification | orchestrator |
| diary | `.pi/journal.md`, `CLAUDE.md` if structural | `scribe` |
| housekeep | stale-branch/PR cleanup, state-file prune | orchestrator |
### Legs that are the orchestrator's alone
The orchestrator (the loop session) delegates every deliberative leg and
keeps three for itself because they are script-shaped and must not be
re-implemented by a model: claim (`issue.sh claim`, which refuses when
someone else holds the issue — the backstop), merge (API calls below), and
housekeep (branch deletion). If a tick does nothing else, it reconciles.
## Model routing
The routing authority is the card's four tiers, reproduced here as the
loop's assignment, not as an argument:
- **T0 `go/mimo-v2.5`** — mechanical gathering, commit/PR/journal prose,
any fan-out. Effectively free; wrong only where wrongness costs a
debugging session, so nothing above takes its word for a *fact*.
- **T1 `qwen/deepseek-v4-pro-0813`** — implement-from-a-written-plan,
understood-diff review, the orchestrator itself. The default session
model; half price 10:0020:00 EDT, which the cron is shaped around.
- **T2 `glm/glm-5.3`** — repo-scale reasoning: selection, planning,
consequences review, validation judgement. Weekly credits with no
rollover: the loop draws them every week by construction, which is the
correct posture. **`glm-5.3-flash`** for anything multimodal
(screenshots, UI inspection).
- **T3 `go/kimi-k3`** — escalation only: two lower tiers already failed,
or the issue is a named gnarly one. A fresh session seeded with the
failing tier's own summary, never a mid-session switch, never parallel,
at most once per day.
The invariant behind all four, from the card: **routing down is cheap,
routing up is expensive.** An implementation that stalls is escalated by
having the T1 session write *what it tried, what failed, what it observed*
and handing that to a new session one tier up. Escalating a session in
place is forbidden in both directions.
Fan-out is allowed on T0 and T1 only (the Go plan's $12/5 h constraint
makes T3 fan-out self-defeating). Critique is the one standing fan-out:
two reviewers, two angles, one synthesis.
## Scheduling
`0 0 10-18 * * 1-5` (local = EDT): hourly on weekdays inside Qwen's
half-price window, clear of the card's ⚠ 26am band (DeepSeek peaks, GLM
loses its off-peak discount — the window the old `yj-backlog` cron sat in,
which this replaces as the loop supersedes it).
- A tick takes a lock (`/tmp/yj-loop.lock`, PID + timestamp). An overrun
tick makes the next fire exit immediately; serialization survives
whatever the scheduler does with overlapping fires.
- ~9 ticks/day; an issue is 25 ticks; **one to two issues per day** is
the natural rate. That also paces the bills without a budget flag.
- The port check is part of reconcile: if `34115` is occupied, the tick
refuses any leg that needs the headless app and defers to the next
tick, without complaint. A human's interactive tier always wins.
## Runtime and ON/OFF
The scheduler (`pi-schedule-prompt`) fires only while a pi session is open
in the job's directory — that limitation is the switch:
- **Worktree:** `git worktree add` a dedicated clone at
`~/.paseo/worktrees/loop/jumpy-hound`. Loop edits happen only there; a
dirty tree there is the loop's business and nobody else's. **Provision
it once before its first push:** `make build-frontend` + `make testdata`
— the pre-push `go-test` hook needs both and refuses without them.
- **Session:** pi in that worktree, `/name loop`. The job is bound to that
session, so another pi elsewhere in the same directory does not
double-fire it.
- **ON:** resume the loop session (`pi --resume loop`) and enable the job.
**OFF:** toggle the job off in `/schedule-prompt`, or close the session.
**Drain** (stop taking new work, finish in flight): set `drain: true` in
the state file.
- **Tune in:** the same `pi --resume loop` — the chat transcript *is* the
loop's log, each tick's reasoning inline, each leg reporting in.
## Identity, claims, and what the loop may touch
The loop operates **as the owner** via `GITEA_TOKEN` (scopes: `read:user`,
`write:issue`, `write:pull`, `write:repository`); pushes ride SSH and need
no token. Every tracker comment the loop writes is prefixed `⟦loop⟧`, so
the collaborator reads it as the pump and not as a person.
It may only ever touch work it created: issues it claimed, branches it
made, PRs it opened. Two mechanisms make that enforced rather than
intentional: `issue.sh claim` refuses an issue somebody else holds, and
reconcile checks `git ls-remote --heads origin` so a branch name collision
from a concurrent session is caught before the first edit.
## Merge lifecycle
- **Only PRs the loop opened.** A collaborator's PR is never merged, never
commented on for pressure, never touched.
- **Every branch is refreshed against main before its merge**, in the
loop worktree — the refresh is where a textual conflict surfaces, as
diff text: hunks the loop authored are resolved there, anything else
is left to a human with a `⟦loop⟧` comment. The protection's
`block_on_outdated_branch` makes the refresh mandatory for adopted
(pre-loop) branches: behind `main`, a PR cannot merge at all.
Required contexts are re-polled on the refreshed head.
- **Merges are one at a time**, each re-reading state — the previous
merge moved `main`, and the next PR's mergeability is recomputed at
its own turn.
- **Post-merge, the `push` run on `main` is watched.** A red main after
a loop merge halts the loop. That run is the only guard against the
class no mergeability check sees: two PRs touching the same file,
merging cleanly, contradicting each other.
- The gate is the protection rule itself, read from the API: contexts
`CI / check*` and `CI / e2e*` green, PR mergeable. (Required approvals
is 0 today; if a second person changes protection rules, the merge
endpoint refuses and the tick stops and reports — human business.)
- `Closes #n` goes **in a commit body inside the branch, one line per
issue, and in the PR body**. Both, because a squash route and a merge
route parse different texts, and this pairing was measured: a comma
list partially matched, five of ten issues.
- After merging: verify against `issue.sh list --state open` that the
issue actually closed; close any straggler naming the merge commit.
`unclaim.yml` strips `Status/In Progress` automatically; it is not
instant, and a re-open does not restore it — the verification is
against the open list, not against the label.
- Merging to `main` fans out to nothing: releases are the manual
`release.yml`, which this loop never runs. The blast radius of a
merge is the main branch's CI, and the critique leg is what stands
before it.
## Verification contract
The tier table is `yellowjacket-dev`'s; the loop re-states nothing above
it except the *division of duty*: the worker runs the tiers the change
demands, and the validator re-reads the issue and checks that the tier
evidence actually answers the claim — a green suite that never touched
the reported surface is a finding, not a pass. Cosmetics are read by a
model that can see (`visual`, the multimodal tier); a change that moves
geometry refreshes its `ui-visual` baseline in the same commit.
`tsc --noEmit` is part of the gate and nothing else runs it. The e2e app
is seeded (`SEED=default`) and stopped after.
## Android / emulator mode
The loop is **device-free by default**: issues whose verification is
physical-device behaviour stay open for humans (the repo's own tags say
which those are). One step of the ladder exists for the rest:
- `{"emulator": true}` in `.pi/loop/state.json` **plus an already-booted
emulator** (`adb devices` answers) opts the loop into building the APK
and using `android-smoke` (crash verification), and `android-screenshot`
/ `android-eval` as rendering evidence for `visual`.
- The loop **never boots or stops an emulator** — that is the user's
machine and their gesture. Boot it with `make android-emulator`
(one-time `make android-setup`, ~3.5 GB, creates the AVD), and
`make android-emulator-stop` when done.
- Real-device-only issues are skipped under either setting.
## Budgets and pacing
Expected spend: dominated by the T1 implementation leg inside the
half-price window (pennies to tens of cents) and T2 on weekly credits;
T3 bounded at one fresh call per day. The card's numbers ($12 per rolling
5 h, $30/week as burst headroom not allowance, GLM reset weekly) are the
sanity cells; the loop's own weekly check compares against them rather
than against the month.
## Cleanup (housekeep leg, once per day)
- Loop-owned branches whose commits are in `origin/main`: deleted, local
and remote.
- Loop-owned PRs open >7 days or red on a second identical CI cause:
commented with what is known (`⟦loop⟧`), and left — never silently
deleted.
- Anything not the loop's (assignee, branch, PR): left exactly as found.
## Pilot phases
- **P0 — supervised.** One tick, user watching the transcript: reconcile,
select, claim, plan. No merge.
- **P1 — observed.** Two ticks ending in the loop's first merge, watched
through CI → merge → verify-close.
- **P2 — unattended.** The schedule left on. Weekly check against the
card's two-minute ritual.
- **Hard stops** (any of these halts the loop and leaves a comment, never
a silent retry): a tick dies twice with no explanation; a merge happens
for a PR the loop did not open; spend outside the cells above by 2×.
## Not now, on purpose
- **Parallel worktrees** — blocked on e2e's exclusive port; viable only
with per-worktree headless ports or CI-only e2e. The shape (
supervisor + per-issue worktrees) is the target, not the first cut.
- **Weekend batch refactors** — DeepSeek off-peak is real but is a
scheduling knob on top of a working pump.
- **More chain files** — the critique fan-out is a chain; the rest stay
orchestrator-legs until two weeks of unattended runs say which legs
are actually fixed-shape.
+108 -1
View File
@@ -6,6 +6,22 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
YellowJacket is a cross-platform desktop music player built with Go (backend) and TypeScript/Lit (frontend), using the Wails framework to bridge them. It supports MP3, FLAC, OGG Vorbis, and WAV playback.
**The three prose documents are split by reader, not by topic** (#50).
`README.md` is the landing page and answers *a user's* questions only —
what it does, which channel installs it on which platform, where its
data lives — with three screenshots in `docs/images/`, captured from the
fixture library (`make sandbox-seed NAME=default``make dev-headless
SEED=default`) so they can be retaken by anyone. `CONTRIBUTING.md` holds
what used to be the second half of that README — prerequisites, the
system libraries, the build and codegen commands, which verification
tier a change demands, the tracker workflow and the commit grammar. This
file stays the deep reference both of them point at, and is the only one
of the three that explains *why* a shape is what it is. A fact that
belongs to a user goes in one place; the packaging channels keep their
own documents (`packaging/*/README.md`, `docs/android-release.md`) and
are linked rather than summarised, because a version-restart note copied
into the README is a second copy to keep true.
## Issues
**The tracker is the source of truth for what is wanted and what is
@@ -122,6 +138,17 @@ has started is a second, staler answer to "what are we doing next".
Numbering is sequential and stable across status moves (a plan keeps
its `NNN-` prefix). Abandoned plans are deleted.
**The autonomous loop** (plan 020, `.pi/skills/yj-loop/`) is the pi
configuration that works the tracker one issue at a time — a cron tick
in a dedicated worktree and session, with the tracker labels as its
state machine. It claims with `issue.sh` like anyone, merges only PRs
it opened once the protection contexts are green, and files what it
finds. Its switch is `.pi/schedule-prompts.json` (gitignored): it runs
only while that pi session is open, and that limitation is the whole
on/off design. Where a loop discovery contradicts this file, this file
is wrong and should be fixed by the diary leg — the loop never quietly
decides otherwise.
## Commands
```bash
@@ -145,7 +172,7 @@ make ui-test # Vitest component/store suite in a real browser (no app)
make ui-visual # Same, including toMatchScreenshot comparisons
make ui-setup # Install the Vitest provider's own Chromium (once)
make bindings-check # Fail if frontend/bindings is stale vs the Go bindings
make skill-check # Fail if .pi/ documents a make target that doesn't exist
make skill-check # Fail if a doc names a make target that doesn't exist
make commit-check # Fail if a commit subject is not a Conventional Commit
make lint # golangci-lint v2 (strict), all three build configurations
make test # All tests with race detector, all three build configurations
@@ -646,6 +673,28 @@ rather than renaming them.
one of the shell's rows, which is what the skip link is absolutely
positioned to avoid.
- `config` — TOML-based settings. Settings page uses HTMX + templ for server-rendered HTML fragments.
**A setter that can reject its argument puts the old value back**, and
that is a correctness rule rather than hygiene (#231). `Save()`
validates the *whole* config, so a value left behind by a failed write
does not merely fail its own call: it fails every later save, of every
unrelated setting — theme, launch page, shortcuts, libraries — for the
rest of the session. Nothing reaches disk, so a restart clears it,
which is exactly what makes the fault invisible and unreportable. One
rejected track-list column list was enough to stop the app saving
anything at all.
Two shapes are safe and a third is the trap. A setter that assigns and
*then* validates snapshots the field first and restores it on the
error path — seven do. `SetLibraryDirectory` is the better shape where
the value can be built on its own: it validates a candidate *before*
assigning, so there is nothing to undo. And a setter whose argument no
validation inspects needs neither — the bools, the favourites playlist
id and the shortcut bindings, plus `SetViewVisible`, which refuses an
unknown, non-hideable or launch-page view up front so
`GeneralConfig.Validate` never sees one it would fail on. Which set a
new setter joins is decided by whether its own `Validate` can reject
it, not by preference.
- `playlist` / `smartplaylist` — Playlist CRUD and rule-based smart playlists.
- `mediacontrols` — OS media controls behind one `Handler`: MPRIS over
D-Bus on desktop Linux, a MediaSession on Android, a no-op stub
@@ -1470,6 +1519,28 @@ live**: a scrim over a menu item is that item's text surface, and the
14px spends its weight below the last legible label, measured at 9.9:1
on the light ramp, whose `bgElevated` is `#e9ecef`.
**And the phone has two sheets, so that rule is one file both read**
(#210). `bottom-nav`'s "More" is capped at the same 85vh and overflows
for the same reason — measured at 424x439 with eight destinations,
`scrollHeight` 412 against `clientHeight` 373, and eleven items at 48px
would be 528, since #25 makes the count the user's. So the two layers
live in `styles/sheet-scroll.css.ts` and each host says only what is
local to it: the colour, handed over as `--yj-sheet-surface` on the same
box, because the nav sheet paints the sidebar's `--yj-bg-surface` and
the context sheet the menus' `--yj-bg-elevated` — a shared rule that
hard-coded either would draw that seam across the other one.
The half that is not the fade is what makes it visible: **nothing inside
the sheet may repaint the surface**, because these are background layers
on the scroller and an opaque child covers them. `menu-surface` already
had it from the other side (`.context-menu-panel[data-sheet]` is
`background-color: transparent`); `app-sidebar`'s host paints
`--yj-bg-surface`, which in the shell is its own background and in the
sheet is a second copy of the sheet's, so `bottom-nav` turns it off.
Measured at 424x439 with the fade adopted and that rule missing: a flat
52,58,64 to the bottom edge with 39px still below, which is the defect
unchanged and every assertion about `background-attachment` passing.
**The playlist submenu is a sheet too, and it had to be.** It is a
`placement="right-start"` flyout, and making the menu full-width moved
its anchor — measured at x 182 to 0, entirely off-screen, so "Add to
@@ -1832,6 +1903,21 @@ is not it.** A `placeholder` is an accname fallback, so an
Explore's search box — the audit's own `a11y.26` — as clean. A sweep
for *empty* names cannot see a *weak* one.
**`title` is the same trap one rung lower, and it defeats the obvious
spec as well as the obvious sweep.** `queue-panel`'s Clear queue and
Add queue to playlist were named by `title` alone, so
`getByRole('button', { name: 'Clear queue' })` matched them **before**
the fix as well as after — a `getByRole` assertion, which is what
catches every other nameless control in this app, would have been
green on the broken build. `title` is the *last* fallback in the
accname order, so content put inside the button later silently
outranks it, and it is the one name a phone cannot show, having no
hover. The property is therefore asserted as *the name is not the
tooltip*: `queue-overlay.spec.ts` removes the `title` attributes and
asks again, which is 1 and 1 with `aria-label` and was measured at 0
and 0 without it. The `title`s stay, because on a desktop they are
also the tooltip for an icon-only control and that is a different job.
**The shell scrolls sideways and not down.** `body` is
`overflow-x: auto; overflow-y: hidden`, and both halves are measured.
Vertically there is nothing to fix: the middle grid row is `1fr` and
@@ -3222,6 +3308,27 @@ its own duplicates apart) — and changing either is invisible against an
existing `YJ_HOME`, whose `config.toml` already holds the old list, so
`make sandbox-seed NAME=default` before believing the app.
**And the *valid* columns are declared twice too, which is the pair
that drifted.** `tracklist.AllColumnIDs` is what the backend accepts;
`COLUMN_DEFS` is what the frontend knows how to draw, and they are not
the same set — `titleArtist` is a definition and not a choice, since it
is the phone's stacked column and is picked by width in
`PHONE_COLUMN_IDS`. Settings built its list from `Object.keys(
COLUMN_DEFS)` and so offered it: **two rows both called "Track Name"**
(#197), the second unselectable, because ticking it sends a column set
Go rejects with `unknown track-list column ID` and `config-page`
swallows that into a `console.error`. `CONFIGURABLE_COLUMN_IDS` is what
the configurator reads now, derived from a `configurable` flag on the
definition, and `settings-column-list.test.ts` reads Go's own list out
of the source rather than writing it down a third time — the rule being
about every column, so checking one checks nothing.
One thing it does **not** fix, because it is reachable from any invalid
input rather than from that row: `SetTrackListColumns` assigns before it
validates, so a rejected list stays in memory and `Save()` validates the
whole config — one tick and **no setting saves for the rest of the
session**, silently. That is #231.
**Event-driven communication**: Backend emits events via Wails runtime; frontend stores subscribe to them. Event names are constants in `backend/events/`.
`frontend/src/events.ts` is **generated** from `backend/events/events.go`
+173
View File
@@ -0,0 +1,173 @@
# Contributing to YellowJacket
This is the contributor's half of the [README](README.md): how to build it, how
to check a change, and how a change gets in. [`CLAUDE.md`](CLAUDE.md) is the
deep reference — the architecture, and the reasons behind the shape of it —
and is worth reading before a change of any size, because most of this
codebase's traps are written down there and nowhere else.
## Building from source
YellowJacket is [Go](https://go.dev/) with a [Lit](https://lit.dev/)/TypeScript
frontend, bridged by [Wails v3](https://wails.io/).
| Tool | Version |
|------|---------|
| Go | 1.25+ |
| Node.js | 22+ |
| pnpm | 10+ |
| Wails CLI | v3 — vendored, no install needed (`go tool wails3`) |
The Wails v3 CLI resolves from the `tool` block in `go.mod`, so there is nothing
to install globally; `make setup` fetches it with the rest of the tooling.
On Linux, install the system libraries Wails needs. v3 builds against GTK4 +
WebKitGTK 6.0 by default:
```bash
sudo apt-get install libasound2-dev libgtk-4-dev libwebkitgtk-6.0-dev # Debian/Ubuntu
sudo pacman -S alsa-lib gtk4 webkitgtk-6.0 # Arch
```
A machine without `webkitgtk-6.0` can still build with `-tags gtk3` against the
older WebKit2GTK 4.1 stack, but that is an escape hatch, not what CI or a
release builds. macOS and Windows need no extra system packages. Run
`go tool wails3 doctor` to check your environment.
```bash
make setup # install tooling, frontend packages and the git hooks
make dev # run with hot-reload
make build-dev # debug build with symbols
make build-prod # production build (stripped and trimmed)
make android # the arm64 APK, into bin/
```
The `Makefile` is the front door and carries a one-line description against
each target; `Taskfile.yml` and `build/<platform>/Taskfile.yml` are the build
implementation behind it and are not called directly.
## Generated code
Two generators run from `go generate ./...`, which `make generate` wraps:
**sqlc** turns `backend/database/sql/queries/` into Go in
`backend/database/sql/sqlcgen/`, and **templ** turns `.templ` files into
`*_templ.go` beside them. Never edit either output by hand — run
`make generate` after touching a `.sql` or a `.templ` file.
The TypeScript bindings in `frontend/bindings/` are generated by `wails3`
rather than by `go generate`, so they are a separate step: `make bindings`
regenerates them and `make bindings-check` fails if they are stale.
`frontend/src/events.ts` is generated too, from `backend/events/events.go`.
A pre-commit hook checks that all of this is fresh, so the usual way to meet it
is a failing commit rather than a bug.
## Checking a change
Run the tier the change actually demands, not the cheapest one.
| Change | Command |
|---|---|
| Go | `make lint` and `make test` — both cover all three build configurations |
| A frontend component or store | `make ui-test` (Vitest in a real Chromium, no backend) |
| A user-visible flow | `make e2e`, against a running `make dev-headless` |
| CSS | `make css-check` — see the Chrome 113 note below |
| Anything cosmetic | look at a screenshot; several bugs here were invisible to every assertion and obvious in an image |
`make test` needs the fixture library, which is generated rather than
committed — it runs `make testdata` itself (about a second).
The end-to-end tier drives the real app with no display at all: `make
dev-headless` starts it in the background on `:34115` (add `SEED=<name>` for a
seeded library, built by `make sandbox-seed NAME=<name>`), `make dev-logs` tails
it and `make dev-stop` stops it. **Check the port before starting one** — if
`:34115` is already answering, someone else's app is there, and a green result
about their build is worse than no result.
Two smaller checks exist because the failure they catch is silent:
`make bindings-check` (stale generated bindings) and `make css-check`, which is
two passes — one fails on a `css` literal ended early by a backtick inside a
comment, the other on a nested CSS rule that begins with a bare element
selector. Chrome 113 is what the reference Android device renders with, and it
drops such a rule without a word.
`make vulncheck` runs govulncheck over the module.
## The issue tracker is the source of truth
Work is described by issues before it is described by branches, and the tracker
is shared with people who cannot see your terminal.
- **Search before starting**, closed issues included: `./scripts/issue.sh search
<terms>`. "That was fixed three weeks ago" is the cheapest possible answer.
- **Claim before the first edit**, not before the commit:
`./scripts/issue.sh claim <n>` sets the assignee, applies `Status/In Progress`
and comments with the branch, so the work is visibly taken *while it is being
done*. It refuses if somebody else holds it — talk to them rather than working
around it.
- **If no issue covers the work, open one first** (`./scripts/issue.sh new`).
- **Findings get filed.** A bug tripped over on the way to something else is an
issue with a reproduction, not a wider diff and not a sentence in a chat log.
- **#73 is the roadmap** and states the order the backlog should be worked in.
`scripts/issue.sh` is the whole interface (`list`, `mine`, `search`, `show`,
`new`, `claim`, `unclaim`, `comment`, `close`, `label`, `depends`, `labels`) and
wants a `GITEA_TOKEN` with `write:issue`. The labels are a taxonomy rather than
tags: `Kind/*`, `Area/*`, `Priority/*`, `Platform/*`, plus `Reviewed/*` and
`Status/*`, of which the last two are exclusive scopes.
## Commits and pull requests
`main` is protected, so a branch and a PR are the only way in. Branch from
`origin/main`, and name the branch after the issue (`fix/140-…`, `feat/25-…`).
Commit subjects are [Conventional Commits](https://www.conventionalcommits.org/)
— `type(scope): subject`, imperative, ≤72 characters — and are enforced by a
`commit-msg` hook and by CI (`make commit-check`). This is load-bearing rather
than decorative: semantic-release reads the **type** to decide the next version,
so a CI-only change is `ci:` and never `fix(ci):`, which would ship a patch
release. `make release-dry` prints what a release would cut right now.
**The closing keyword goes in the commit body**, one issue per line, because
Gitea parses commit messages that reach `main` and does not parse the PR body:
```
docs: rewrite the README as a landing page
<why>
Closes #50
```
A PR body carries a commit-to-issue table, the verification you actually ran
(with results), and a `Closes` list for whoever reads it.
## Style
- **Go** — golangci-lint v2, strict: `err113` (static errors), `nlreturn`,
`wsl_v5`, `godot`, `sloglint`, `perfsprint`, and imports grouped stdlib →
third-party → `yellowjacket/…` by gci.
- **TypeScript** — strict mode, no implicit `any`, no unused locals or
parameters.
- Match the surrounding code. Where `CLAUDE.md` explains why something is shaped
the way it is, that shape is load-bearing and there is usually a test pinning
it.
Hooks do most of the enforcing (`lefthook.yml`, installed by `make setup`):
pre-commit runs vet, lint, the codegen checks, the frontend typecheck and the
two CSS checks in parallel; pre-push runs the Go suite and the UI tier,
deliberately one after the other rather than together.
## Where the rest of the documentation is
- [`CLAUDE.md`](CLAUDE.md) — architecture and constraints, in depth.
- [`docs/PROFILING.md`](docs/PROFILING.md) — Go pprof and frontend profiling.
- [`docs/android-release.md`](docs/android-release.md) — the APK, its signing
key, and what the release workflow checks.
- [`docs/index-cache.md`](docs/index-cache.md) — the search-index build cache
and why it has a snapshot.
- [`packaging/arch/README.md`](packaging/arch/README.md),
[`packaging/homebrew/README.md`](packaging/homebrew/README.md) — the two
package channels.
- `.planning/` — design documents and measured history, not a queue. The queue
is the tracker.
+1 -1
View File
@@ -192,7 +192,7 @@ css-check: ## Fail on a css`` literal ended early by a backtick, or a nested rul
# Every command in them is a make target on purpose, so this is
# checkable. It also asserts AGENTS.md is a symlink to CLAUDE.md, so the
# two harnesses cannot drift onto two descriptions of one project.
skill-check: ## Fail if the agent docs name a missing make target, or AGENTS.md is not a symlink
skill-check: ## Fail if the docs name a missing make target, or AGENTS.md is not a symlink
@./scripts/skill-check.sh
# Conventional Commits, which CLAUDE.md claimed CI enforced for a long
+107 -80
View File
@@ -2,112 +2,139 @@
*Music how it was meant to bee.*
YellowJacket is a fast, cross-platform desktop music player for your local
collection. It plays your files, keeps your library tidy, and helps you discover
and organize your music — all in a clean, responsive interface. No accounts, no
streaming, no telemetry: just your music on your machine.
YellowJacket plays the music you already own. Point it at your folders and it
scans them, reads the tags and the cover art, and gives you a library you can
browse, search, queue and tidy up — on your own machine, with no account, no
streaming service and no telemetry.
Runs on **Linux**, **macOS**, and **Windows**.
It plays **MP3**, **FLAC**, **OGG Vorbis** and **WAV**, on **Linux** and
**Android**, and builds from source on **macOS**.
## Features
![The track list, with something playing](docs/images/library.png)
### Play your music
- Plays **MP3, FLAC, OGG Vorbis, and WAV**
- Play, pause, seek, and volume control with a mute toggle
- Gapless, glitch-free seeking backed by a read-ahead buffer
- A queue you can add to, reorder, and shuffle, with play-next support
- Shuffle and repeat (off / all / one)
- Picks up right where you left off — remembers your track, position, and volume between sessions
- Media-key and MPRIS support on Linux, so your desktop's playback controls just work
## What it does
### Keep your library organized
- Point it at your music folders and it scans them automatically
- Reads tags and embedded cover art, and de-duplicates artwork so it isn't stored twice
- Incremental sync — only new or changed files get reprocessed, and deleted files are cleaned up
- Browse by **album**, **artist**, or **genre**, or search across everything
- Mark favorites and see what you've been listening to with play history
- Edit track tags directly when something's off
**Plays your files.** Play, pause, seek and volume with a mute toggle; a
read-ahead buffer so seeking is instant rather than gappy; a queue you can add
to, reorder and shuffle, with play-next; shuffle and repeat (off / all / one).
It remembers the track, the position and the queue between sessions, and it
answers your desktop's media keys — MPRIS on Linux, a media notification and
lock-screen controls on Android.
### Playlists
- Create playlists, drag tracks in, and reorder them
- **Smart playlists** that build themselves from rules (by genre, rating, play count, and more)
- Pin a default playlist and spot duplicate tracks at a glance
**Keeps the library tidy.** It scans the folders you give it and rescans only
what changed, so a big library costs its full scan once. It de-duplicates
embedded cover art rather than storing the same image a hundred times, notices
files that have gone away, and spots duplicate tracks. Browse by album, artist
or genre, search across everything, mark favourites, and see what you have been
playing.
### Discover and clean up (powered by MusicBrainz)
- **Explore** — browse artists, releases, and genres from the MusicBrainz catalog, not just what's already in your library
- **Auto-tag** — match your files against MusicBrainz to fill in correct artist, album, and track metadata, with a review step before anything is written
- **Lyrics search** — find a track by a line you remember
**Playlists, and playlists that write themselves.** Drag tracks in and reorder
them, or describe what you want — genre, play count, how long since you played
it — and let a smart playlist keep itself up to date.
**Explore and auto-tag, from the MusicBrainz catalog.** Explore browses artists,
releases and genres from the catalog rather than only from what you own, so an
album page can tell you that you have nine of its twelve tracks. Auto-tag
matches your files against MusicBrainz and fills in the metadata that is
missing, with a review step before anything is written to disk. Lyrics search
finds a track from a line you remember.
Explore needs its catalog, which is a one-off ~0.6 GB download from
**Settings → Search Index**. It asks first on a metered connection, and
everything else in the app works without it.
## Install
Download the latest build for your platform from the
Every download comes from the
[releases page](https://git.ljones.me/yonlu/yellowjacket/releases).
| Platform | Download |
|----------|----------|
| Linux | `yellowjacket-linux-amd64` |
| macOS | `yellowjacket-darwin-universal.app.zip` (Apple Silicon + Intel) |
| Windows | `yellowjacket-windows-amd64.exe` |
### Linux
Prefer to build it yourself? See [Building from source](#building-from-source).
Download `yellowjacket-<version>-linux-amd64.tar.gz` from the latest release and
unpack it. It holds the binary, a `.desktop` entry and an icon.
## Getting started
On **Arch**, install it from the package registry instead and get updates with
the rest of your system — the one-time key import and `pacman.conf` block are in
[`packaging/arch/README.md`](packaging/arch/README.md):
```bash
sudo pacman -Sy yellowjacket
```
### Android
Install the APK from the release page, or from the URL below, which always
points at the newest build:
```
https://git.ljones.me/api/packages/yonlu/generic/yellowjacket-android/latest/yellowjacket.apk
```
That URL needs no credentials, so [Obtainium](https://obtainium.imranr.dev/) can
poll it directly and keep the app up to date. The build is `arm64-v8a` only, and
[`docs/android-release.md`](docs/android-release.md) says why.
### macOS
Homebrew builds it from source on your own Mac — there is no prebuilt `.app`,
because a signed macOS bundle needs a macOS machine to produce it and the
release runner is a Linux container.
```bash
brew install shadow-puppet/yellowjacket/yellowjacket
```
See [`packaging/homebrew/README.md`](packaging/homebrew/README.md).
### Windows
Not published. It cross-compiles cleanly, but no Windows build of this app has
ever been *run*, and nothing here can exercise one — so shipping it would be a
promise that cannot be kept. You can still build it yourself: see
[`CONTRIBUTING.md`](CONTRIBUTING.md).
### Coming from a 1.x install?
Versions restarted at **0.0.1** when releases became automatic, which every
package manager reads as a downgrade. It costs one reinstall, once — the details
are with each channel: [Homebrew](packaging/homebrew/README.md#upgrading-from-1x-needs-a-reinstall-once),
[Android](docs/android-release.md#the-1x-installs-cannot-be-upgraded-to-00x).
## First run
1. Launch YellowJacket.
2. Open **Settings** and add the folder(s) where your music lives.
3. Let the initial scan finish — you'll see progress as it works.
4. Browse by album, artist, or genre, queue something up, and press play.
2. Add the folder your music lives in — the first-run wizard asks, and
**Settings → Libraries** is where you add more later.
3. Watch the scan finish. It reports progress, and you can browse while it runs.
4. Queue something and press play.
Your library and settings are stored locally:
Your library and settings stay on your machine:
| | Linux / macOS | Windows |
|---|---|---|
| Config | `~/.config/yellowjacket/` | `%LOCALAPPDATA%\yellowjacket\config` |
| Library data | `~/.local/share/yellowjacket/` | `%LOCALAPPDATA%\yellowjacket\data` |
## Building from source
Setting `YJ_HOME` moves both, which is how you keep a second library separate.
YellowJacket is built with [Go](https://go.dev/) and a
[Lit](https://lit.dev/)/TypeScript frontend, bridged by the
[Wails](https://wails.io/) framework.
## More screenshots
**Prerequisites**
An album page knows what you own, and says so:
| Tool | Version |
|------|---------|
| Go | 1.25+ |
| Node.js | 22+ |
| pnpm | 10+ |
| Wails CLI | v3 — vendored, no install needed (`go tool wails3`) |
![An album page, with two discs and the transport playing](docs/images/album.png)
The Wails v3 CLI resolves from the `tool` block in `go.mod`, so there is nothing
to install globally; `make setup` fetches it with the rest of the tooling.
The home page suggests somewhere to start rather than opening on a wall of
everything:
On Linux, install the system libraries Wails needs. v3 builds against GTK4 +
WebKitGTK 6.0 by default:
![The home page's shelves](docs/images/home.png)
```bash
sudo apt-get install libasound2-dev libgtk-4-dev libwebkitgtk-6.0-dev # Debian/Ubuntu
sudo pacman -S alsa-lib gtk4 webkitgtk-6.0 # Arch
```
## Contributing, and the rest of the documentation
A machine without `webkitgtk-6.0` can still build with `-tags gtk3` against the
older WebKit2GTK 4.1 stack, but that is an escape hatch, not what CI or a
release builds.
macOS and Windows need no extra system packages. Run `go tool wails3 doctor` to
check your environment.
**Build**
```bash
make setup # install tooling and git hooks
make dev # run with hot-reload
make build-prod # produce a release binary
```
More detail for contributors lives in [`CLAUDE.md`](./CLAUDE.md) — the
architecture, the conventions and the reasons behind them. What is
being worked on is [the issue
tracker](https://git.ljones.me/yonlu/yellowjacket/issues); #73 is the
roadmap.
- [`CONTRIBUTING.md`](CONTRIBUTING.md) — build it from source, run the tests,
and how a change gets in.
- [`CLAUDE.md`](CLAUDE.md) — the deep reference: the architecture and the reasons
behind the shape of it.
- [The issue tracker](https://git.ljones.me/yonlu/yellowjacket/issues) is what
is wanted and what is being worked on; **#73** is the roadmap.
- [Releases](https://git.ljones.me/yonlu/yellowjacket/releases) double as the
changelog — every one is generated from the commits it contains.
+39
View File
@@ -303,6 +303,24 @@ func (c *Config) GetLibraryDirectory() string {
return string(c.Library.DirectoryPath)
}
// A rejected setter puts the old value back, and that is not tidiness
// (#231). Save validates the *whole* config, so a value left behind by
// a failed write does not merely fail its own call: it fails every
// later save, of every unrelated setting, silently and for the rest of
// the session. Nothing reaches disk, so a restart clears it -- which
// is exactly what makes the fault hard to see and impossible to report.
//
// The setters below that assign and then validate therefore snapshot
// the field first and restore it on the error path. SetLibraryDirectory
// is the other safe shape and the better one where the value can be
// built on its own: it validates a candidate *before* assigning
// anything, so there is nothing to undo.
//
// Not every setter needs either. A bool, an int64 and the shortcut
// bindings pass through no validation that can reject them, and
// SetViewVisible refuses an unknown, non-hideable or launch-page view
// up front, so GeneralConfig.Validate never sees one it would fail on.
// SetLibraryDirectory validates and saves a new library directory,
// then emits the LibraryConfigChanged event so listeners (e.g. the
// Library scanner) can react.
@@ -360,11 +378,14 @@ func (c *Config) SetScanConcurrency(mode string) error {
c.Library.ApplyDefaults()
}
previous := c.Library.ScanConcurrency
c.Library.ScanConcurrency = library.ScanConcurrency(
mode,
)
if err := c.Library.Validate(); err != nil {
c.Library.ScanConcurrency = previous
return fmt.Errorf(
"invalid scan concurrency mode: %w", err,
)
@@ -455,9 +476,12 @@ func (c *Config) SetThemeAccentColor(
c.Theme.ApplyDefaults()
}
previous := c.Theme.AccentColor
c.Theme.AccentColor = color
if err := c.Theme.Validate(); err != nil {
c.Theme.AccentColor = previous
return fmt.Errorf(
"invalid theme accent color: %w", err,
)
@@ -488,9 +512,12 @@ func (c *Config) SetThemeBackgroundShade(
c.Theme.ApplyDefaults()
}
previous := c.Theme.BackgroundShade
c.Theme.BackgroundShade = theme.BackgroundShade(shade)
if err := c.Theme.Validate(); err != nil {
c.Theme.BackgroundShade = previous
return fmt.Errorf(
"invalid theme background shade: %w", err,
)
@@ -544,9 +571,12 @@ func (c *Config) SetDefaultPage(page string) error {
c.General.ApplyDefaults()
}
previous := c.General.DefaultPage
c.General.DefaultPage = View(page)
if err := c.General.Validate(); err != nil {
c.General.DefaultPage = previous
return fmt.Errorf(
"invalid default page: %w", err,
)
@@ -591,9 +621,12 @@ func (c *Config) SetQueueFallback(mode string) error {
c.General.ApplyDefaults()
}
previous := c.General.QueueFallback
c.General.QueueFallback = QueueFallback(mode)
if err := c.General.Validate(); err != nil {
c.General.QueueFallback = previous
return fmt.Errorf(
"invalid queue fallback: %w", err,
)
@@ -801,9 +834,12 @@ func (c *Config) SetTrackListColumns(
c.TrackList = &tracklist.Config{}
}
previous := c.TrackList.Columns
c.TrackList.Columns = columns
if err := c.TrackList.Validate(); err != nil {
c.TrackList.Columns = previous
return fmt.Errorf(
"invalid track-list columns: %w", err,
)
@@ -901,9 +937,12 @@ func (c *Config) SetFavoritesIconStyle(
c.Favorites.ApplyDefaults()
}
previous := c.Favorites.IconStyle
c.Favorites.IconStyle = favorites.IconStyle(style)
if err := c.Favorites.Validate(); err != nil {
c.Favorites.IconStyle = previous
return fmt.Errorf(
"invalid favorites icon style: %w", err,
)
+262
View File
@@ -0,0 +1,262 @@
package config
import (
"log/slog"
"path/filepath"
"testing"
"yellowjacket/backend/library"
"yellowjacket/backend/tracklist"
)
// newSavableConfig builds a loaded, valid config in a temp directory,
// so Save() writes rather than refusing with errSaveBeforeLoad.
//
// The library directory is real and set, because Config.Validate only
// validates the Library section when DirectoryPath is non-empty -- an
// empty one would hide a poisoned ScanConcurrency from the whole-config
// save that is the symptom under test.
func newSavableConfig(t *testing.T) *Config {
t.Helper()
c := &Config{
logger: slog.Default(),
filePath: filepath.Join(t.TempDir(), "config.toml"),
Library: &library.Config{
DirectoryPath: library.Directory(t.TempDir()),
},
}
c.applyDefaults()
if err := c.Load(); err != nil {
t.Fatalf("Load() error: %v", err)
}
if err := c.Save(); err != nil {
t.Fatalf("Save() on a fresh config error: %v", err)
}
return c
}
// TestSetterRejectionDoesNotPoisonTheConfig is the whole of #231.
//
// Every setter here assigns to the in-memory config and then validates.
// When the validation rejects the argument, the rejected value has to go
// back -- not because the caller sees it (it gets an error either way),
// but because Config.Save() validates the *whole* config. A value left
// behind by a failed setter therefore fails every later save, of every
// unrelated setting, silently and for the rest of the session.
//
// So each case asserts three things in order: the setter reports the
// error, the getter still reports the old value, and an unrelated save
// still works. The third is the one the user feels.
func TestSetterRejectionDoesNotPoisonTheConfig(t *testing.T) {
t.Parallel()
cases := []struct {
name string
// reject calls the setter with an argument its own Validate
// refuses.
reject func(*Config) error
// read reports the value the setter writes, so the rollback is
// asserted on the config rather than only on the save.
read func(*Config) string
}{
{
name: "scan concurrency",
reject: func(c *Config) error {
return c.SetScanConcurrency("telepathy")
},
read: (*Config).GetScanConcurrency,
},
{
name: "theme accent colour",
reject: func(c *Config) error {
return c.SetThemeAccentColor("not-a-hex")
},
read: (*Config).GetThemeAccentColor,
},
{
name: "theme background shade",
reject: func(c *Config) error {
return c.SetThemeBackgroundShade("chartreuse")
},
read: (*Config).GetThemeBackgroundShade,
},
{
name: "default page",
reject: func(c *Config) error {
return c.SetDefaultPage("nowhere")
},
read: (*Config).GetDefaultPage,
},
{
name: "queue fallback",
reject: func(c *Config) error {
return c.SetQueueFallback("improvise")
},
read: (*Config).GetQueueFallback,
},
{
name: "favorites icon style",
reject: func(c *Config) error {
return c.SetFavoritesIconStyle("asterisk")
},
read: (*Config).GetFavoritesIconStyle,
},
{
name: "track-list columns",
reject: func(c *Config) error {
// titleArtist is a drawing definition, not a
// configurable column (#197), so it is exactly what
// the frontend used to be able to send.
return c.SetTrackListColumns([]tracklist.Column{
{ID: "titleArtist"},
})
},
read: func(c *Config) string {
return columnIDs(c.GetTrackListColumns())
},
},
{
name: "track-list columns, duplicated",
reject: func(c *Config) error {
// The route #197 closed was one invalid id; a
// duplicate is the one still reachable from a client
// that assembles the list itself.
return c.SetTrackListColumns([]tracklist.Column{
{ID: tracklist.ColTrackName},
{ID: tracklist.ColTrackName},
})
},
read: func(c *Config) string {
return columnIDs(c.GetTrackListColumns())
},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
c := newSavableConfig(t)
before := tc.read(c)
if err := tc.reject(c); err == nil {
t.Fatal("setter accepted an invalid value, want an error")
}
if after := tc.read(c); after != before {
t.Errorf(
"value after a rejected write = %q, want the previous %q",
after, before,
)
}
// The symptom: an unrelated setting can no longer be saved.
if err := c.SetPopupVolume(true); err != nil {
t.Errorf("an unrelated setter failed after a rejected write: %v", err)
}
if err := c.Save(); err != nil {
t.Errorf("Save() failed after a rejected write: %v", err)
}
})
}
}
// TestRejectedSetterLeavesNothingOnDisk pairs with the sweep above: the
// rollback must not be undone by what the file already holds, so a
// config reloaded from disk after a rejected write agrees with memory.
func TestRejectedSetterLeavesNothingOnDisk(t *testing.T) {
t.Parallel()
c := newSavableConfig(t)
if err := c.SetThemeAccentColor("#123456"); err != nil {
t.Fatalf("SetThemeAccentColor() error: %v", err)
}
if err := c.SetThemeAccentColor("not-a-hex"); err == nil {
t.Fatal("SetThemeAccentColor accepted a non-colour, want an error")
}
reloaded := &Config{logger: slog.Default(), filePath: c.filePath}
reloaded.applyDefaults()
if err := reloaded.Load(); err != nil {
t.Fatalf("Load() error: %v", err)
}
if got := reloaded.GetThemeAccentColor(); got != "#123456" {
t.Errorf("accent colour on disk = %q, want %q", got, "#123456")
}
if c.GetThemeAccentColor() != reloaded.GetThemeAccentColor() {
t.Errorf(
"in-memory accent %q disagrees with disk %q after a rejected write",
c.GetThemeAccentColor(), reloaded.GetThemeAccentColor(),
)
}
}
// TestSetLibraryDirectoryValidatesBeforeAssigning pins the precedent the
// seven rolled-back setters follow: this one has always built and
// validated a candidate before assigning, so a bad path never reaches
// the config at all.
func TestSetLibraryDirectoryValidatesBeforeAssigning(t *testing.T) {
t.Parallel()
c := newSavableConfig(t)
before := c.GetLibraryDirectory()
if err := c.SetLibraryDirectory(filepath.Join(t.TempDir(), "no-such-dir")); err == nil {
t.Fatal("SetLibraryDirectory accepted a missing directory, want an error")
}
if after := c.GetLibraryDirectory(); after != before {
t.Errorf("library directory = %q, want the previous %q", after, before)
}
if err := c.Save(); err != nil {
t.Errorf("Save() failed after a rejected library directory: %v", err)
}
}
// TestSetViewVisibleRefusesBeforeAssigning covers the other setter left
// out of the rollback pass: it guards its own argument up front, so
// GeneralConfig.Validate never sees a view it would reject.
func TestSetViewVisibleRefusesBeforeAssigning(t *testing.T) {
t.Parallel()
c := newSavableConfig(t)
if err := c.SetViewVisible("no-such-view", false); err == nil {
t.Fatal("SetViewVisible accepted an unknown view, want an error")
}
if err := c.SetViewVisible(c.GetDefaultPage(), false); err == nil {
t.Fatal("SetViewVisible hid the launch page, want an error")
}
if err := c.Save(); err != nil {
t.Errorf("Save() failed after a refused view visibility change: %v", err)
}
}
// columnIDs renders a column list for comparison in the table above.
func columnIDs(cols []tracklist.Column) string {
ids := make([]byte, 0, len(cols)*8)
for i, col := range cols {
if i > 0 {
ids = append(ids, ',')
}
ids = append(ids, col.ID...)
}
return string(ids)
}
+21 -3
View File
@@ -7,6 +7,7 @@ import (
"path/filepath"
"runtime"
"strings"
"syscall"
"testing"
)
@@ -17,6 +18,21 @@ import (
// stubYtDlp writes an executable script that echoes the given stdout
// and returns it as a provider config binary path.
//
// The write is held under syscall.ForkLock, and that is not tidiness:
// the kernel refuses to exec a file that is open for writing anywhere
// in the process, and these tests are parallel, so a *sibling* test's
// fork can duplicate this descriptor in the moment it is open and
// carry it past our close — the exec a moment later then fails with
// ETXTBSY, "text file busy". That is #146, seen once in CI and once
// locally, on trees containing no Go at all. Closing sooner is not
// available (os.WriteFile has already closed the file before anything
// execs it) and O_CLOEXEC does not help, because the window is between
// another goroutine's fork and its own exec. ForkLock is the lock
// syscall.forkExec takes across that fork, so holding it here means no
// child can exist while the descriptor does. Measured on this helper
// under 12 concurrent writers: 176-189 of 2400 execs refused without
// it, 0 of 2400 with it.
func stubYtDlp(t *testing.T, script string) string {
t.Helper()
@@ -26,9 +42,11 @@ func stubYtDlp(t *testing.T, script string) string {
path := filepath.Join(t.TempDir(), "yt-dlp")
if err := os.WriteFile(
path, []byte("#!/bin/sh\n"+script), 0o700,
); err != nil {
syscall.ForkLock.Lock()
err := os.WriteFile(path, []byte("#!/bin/sh\n"+script), 0o700)
syscall.ForkLock.Unlock()
if err != nil {
t.Fatalf("write stub: %v", err)
}
+6 -3
View File
@@ -63,12 +63,15 @@ func Parse(r io.Reader) ([]Chunk, error) {
return nil, err
}
data := make([]byte, size)
if _, err := io.ReadFull(r, data); err != nil {
// Copied rather than allocated up front, as ID3Chunk does: the
// size is four bytes off the file, so a truncated one is free to
// declare a chunk larger than the whole of itself.
var data bytes.Buffer
if _, err := io.CopyN(&data, r, int64(size)); err != nil {
return nil, fmt.Errorf("read chunk data for %q: %w", id, err)
}
chunks = append(chunks, Chunk{ID: id, Data: data})
chunks = append(chunks, Chunk{ID: id, Data: data.Bytes()})
// Odd-length chunks have a padding byte. Lenient: if the
// read fails (e.g. EOF), just break rather than error.
+43
View File
@@ -4,6 +4,7 @@ import (
"bytes"
"encoding/binary"
"errors"
"runtime"
"testing"
"yellowjacket/backend/riff"
@@ -209,3 +210,45 @@ func TestParse_ReadsEveryChunkInOrder(t *testing.T) {
t.Errorf("odd chunk data: got %q, want %q", chunks[1].Data, "INFOodd")
}
}
// A chunk size is four bytes read off the file, so a truncated or
// malformed WAV is free to declare a chunk larger than the whole of
// itself. Parse must grow with what arrives rather than with what was
// claimed.
//
// This measures the allocation instead of the error because the error
// is the same either way: a build sizing its buffer from the header
// reports the truncation correctly, having asked the allocator for a
// gigabyte on the way. Deliberately not parallel — TotalAlloc is
// process-wide, and a test paused beside another one is measuring it
// too.
func TestParse_DoesNotAllocateWhatAChunkClaims(t *testing.T) {
// Large enough that a header-sized buffer is unmistakable, in a
// container of a few dozen bytes.
const declared = 1 << 30
var raw bytes.Buffer
raw.WriteString("RIFF")
_ = binary.Write(&raw, binary.LittleEndian, uint32(declared+12))
raw.WriteString("WAVE")
raw.WriteString("data")
_ = binary.Write(&raw, binary.LittleEndian, uint32(declared))
raw.WriteString("and then the file ends")
var before, after runtime.MemStats
runtime.GC()
runtime.ReadMemStats(&before)
if _, err := riff.Parse(bytes.NewReader(raw.Bytes())); err == nil {
t.Fatal("Parse: got nil error for a chunk larger than the file holding it")
}
runtime.ReadMemStats(&after)
if grew := after.TotalAlloc - before.TotalAlloc; grew > 1<<20 {
t.Errorf("Parse allocated %d bytes reading a %d-byte file whose chunk header claimed %d",
grew, raw.Len(), declared)
}
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 83 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 128 KiB

+56
View File
@@ -157,6 +157,62 @@ test.describe('an overlaid queue says it is over the content', () => {
});
});
/**
* #170 — the other two buttons in that same row.
*
* Clear queue and Add queue to playlist predate the close button and
* were named by a `title` attribute and nothing else. Unlike the
* sliders in `control-names.spec.ts`, that is not a *missing* name:
* `title` is the last fallback in the accname order, so
* `getByRole('button', { name: 'Clear queue' })` matched them before
* this fix as well as after it — measured, 1 and 1. A sweep for empty
* names cannot see a weak one, which is `a11y.26`'s complaint and the
* reason this file could have grown a green test that proved nothing.
*
* So the name is asserted twice, and the second assertion is the one
* that fails on the broken build. Taking the tooltip away and asking
* again is the property in words: **the name is not the tooltip**. It
* is what makes the button survive content being put inside it later,
* and it is the only one of the two a phone has — there is no hover on
* the surface #55 turned into a full screen. Measured on `main` before
* the fix: 0 and 0.
*
* Both buttons are disabled here, because the queue starts empty and
* naming is not enablement. A disabled button is still in the
* accessibility tree, which is exactly where the complaint was.
*/
test.describe('the queue header says what its actions do', () => {
const ACTIONS = ['Clear queue', 'Add queue to playlist'];
test('names both of the older actions', async ({ app }) => {
await openQueue(app);
for (const name of ACTIONS) {
await expect(
app.getByRole('button', { name, exact: true }),
).toHaveCount(1);
}
});
test('and the names do not come from the tooltip', async ({ app }) => {
await openQueue(app);
await app.locator('#queue-panel').evaluate((el) => {
for (const button of el.shadowRoot!.querySelectorAll(
'.header-action-button',
)) {
button.removeAttribute('title');
}
});
for (const name of ACTIONS) {
await expect(
app.getByRole('button', { name, exact: true }),
).toHaveCount(1);
}
});
});
/**
* The inline panel is the mode that already worked, and the one every
* other queue spec is written against. It keeps its resize handle and
+11 -6
View File
@@ -103,12 +103,17 @@ async function queueSixAndOpen(app: Page): Promise<void> {
*
* `explore-link` routes a track name to its *album's* page, so a
* track with no album renders a name that navigates nowhere — and
* the fixture library deliberately contains two (`01 Tone A`,
* `02 Tone B`). Which tracks arrive first is `audio_files.id`
* order, i.e. the order the **scan** inserted them, which depends
* on concurrency and directory traversal: locally the first eight
* all had albums and the spec passed twice over, and CI rebuilds
* its seed with a real scan and got a different eight.
* the fixture library deliberately contains two,
* `unsorted/no-tags-at-all.mp3` and `unsorted/title-only.mp3`.
* (It contained four until #104: the two WAVs under `Field
* Recordings/Test Tones` had been tagged on disk all along and
* scan in with their album now, so they are ordinary tracks and
* not examples of this.) Which tracks arrive first is
* `audio_files.id` order, i.e. the order the **scan** inserted
* them, which depends on concurrency and directory traversal:
* locally the first eight all had albums and the spec passed twice
* over, and CI rebuilds its seed with a real scan and got a
* different eight.
*
* Asking for what the test needs is the fix. It is not a
* narrowing: every assertion here wants an ordinary track, and
@@ -4,6 +4,7 @@ import '@awesome.me/webawesome/dist/components/icon/icon.js';
import '@awesome.me/webawesome/dist/components/drawer/drawer.js';
import type WaDrawer from '@awesome.me/webawesome/dist/components/drawer/drawer.js';
import { designTokens } from '../../styles/tokens.css';
import { sheetScrollFade } from '../../styles/sheet-scroll.css';
import '../sidebar/app-sidebar.js';
import { nameDialog } from '@utils/name-dialog';
import { ICON_PLAYLIST } from '@utils/icon-language';
@@ -167,6 +168,15 @@ export class BottomNav extends LitElement {
overflow: hidden;
}
/* And this list does not fit (#210): measured at 424x439 with
the seed's eight destinations, the body is scrollHeight 412
against clientHeight 373, and eleven items at 48px would be
528 -- the count is the user's since #25. So the sheet says
where the fold is, with styles/sheet-scroll.css's two layers
rather than a second answer to the question #207 settled for
the context sheet. The colour is the local half: the sidebar
paints --yj-bg-surface, so the cover does too, or the fade
draws the menus' grey across the bottom of this one. */
wa-drawer::part(body) {
padding: 0;
/* A scroll that reaches the end of this list must not
@@ -177,6 +187,23 @@ export class BottomNav extends LitElement {
on a gesture-navigation phone -- the same allowance the
bar itself makes above. */
padding-bottom: env(safe-area-inset-bottom, 0);
--yj-sheet-surface: var(--yj-bg-surface, #212529);
${sheetScrollFade}
}
/* And the sheet paints that surface once. The sidebar's host
paints the same grey -- which in the shell is the sidebar's
own background and here is a second, opaque copy of the
sheet's, drawn *over* the body's layers. So the fade was
painted and then covered: measured at 424x439 before this
rule, the last 32px read a flat 52,58,64 with 39px still
below. menu-surface meets the same requirement from the
other side, where .context-menu-panel[data-sheet] is
background-color: transparent; nothing changes visually
here, because the colour underneath is the one being
removed. */
app-sidebar {
background-color: transparent;
}
/* A sheet is dragged at with a thumb, so it says where its top
@@ -51,7 +51,7 @@ import type { BackgroundShade } from '@store/theme-store';
import type { IconStyle } from '@store/favorites-store';
import {
COLUMN_DEFS,
ALL_COLUMN_IDS,
CONFIGURABLE_COLUMN_IDS,
} from '@components/track-list/columns';
import './config-field';
@@ -1640,7 +1640,7 @@ export class ConfigPage extends ViewLifecycleMixin(LitElement) {
...this.trackListCtrl.columnIds,
];
const disabledIds = ALL_COLUMN_IDS.filter(
const disabledIds = CONFIGURABLE_COLUMN_IDS.filter(
(id) => !enabledIds.includes(id),
);
@@ -2,12 +2,14 @@ import { LitElement, html, css, nothing } from 'lit';
import { customElement, state, query } from 'lit/decorators.js';
import '@awesome.me/webawesome/dist/components/dialog/dialog.js';
import '@awesome.me/webawesome/dist/components/icon/icon.js';
import { EventsOn } from '@runtime/runtime';
import {
AddLibrary,
GetAllLibrariesWithTrackCounts,
} from '@go/library/library.js';
import { describeError, explainError } from '@utils/describe-error';
import { nameDialogsIn } from '@utils/name-dialog';
import { Events } from '../../events';
import { pickDirectory } from '../../utils/pick-directory';
/**
@@ -19,6 +21,13 @@ import { pickDirectory } from '../../utils/pick-directory';
* prompting the user to pick their music folder, registers it through the
* library CRUD API, and dismisses itself. AddLibrary emits LibraryAdded
* and kicks off the initial scan automatically.
*
* **The dismissal follows the library existing, not the button being
* pressed.** `AddLibrary` emits `LibraryAdded` whoever calls it, so the
* wizard waits on the state it exists to wait for rather than on a step
* in its own flow — a library arriving by any other route (Settings, a
* direct call) leaves a full-screen modal up otherwise, intercepting
* every pointer event.
*/
@customElement('first-run-wizard')
export class FirstRunWizard extends LitElement {
@@ -37,9 +46,18 @@ export class FirstRunWizard extends LitElement {
/** Error message from a failed pick/save, if any. */
@state() private errorMessage = '';
/** Unsubscribe from LibraryAdded, while this element is connected. */
private cancelLibraryAdded?: () => void;
override async connectedCallback(): Promise<void> {
super.connectedCallback();
// Subscribed before the read below, so a library arriving while
// that call is in flight is not answered with a stale empty list.
this.cancelLibraryAdded = EventsOn(Events.LibraryAdded, () => {
this.dismiss();
});
try {
const existing = await GetAllLibrariesWithTrackCounts();
@@ -54,6 +72,8 @@ export class FirstRunWizard extends LitElement {
return;
}
if (this.finished) return;
this.active = true;
await this.updateComplete;
@@ -61,6 +81,13 @@ export class FirstRunWizard extends LitElement {
if (this.dialog) this.dialog.open = true;
}
override disconnectedCallback(): void {
this.cancelLibraryAdded?.();
this.cancelLibraryAdded = undefined;
super.disconnectedCallback();
}
static override styles = css`
wa-dialog {
--width: 480px;
@@ -239,6 +266,20 @@ export class FirstRunWizard extends LitElement {
if (!this.finished) e.preventDefault();
};
/**
* Close, and stay closed: a library exists, so setup is over.
*
* `finished` is set first, or `preventClose` cancels the hide this
* asks for.
*/
private dismiss(): void {
this.finished = true;
if (this.dialog) this.dialog.open = false;
this.active = false;
}
private handleChoose = async (): Promise<void> => {
this.errorMessage = '';
@@ -264,11 +305,7 @@ export class FirstRunWizard extends LitElement {
try {
await AddLibrary(this.selectedDirectory);
this.finished = true;
if (this.dialog) this.dialog.open = false;
this.active = false;
this.dismiss();
} catch (err) {
this.errorMessage = explainError(
err,
@@ -65,6 +65,7 @@ import '@awesome.me/webawesome/dist/components/popup/popup.js';
import '@awesome.me/webawesome/dist/components/dialog/dialog.js';
import type WaPopup from '@awesome.me/webawesome/dist/components/popup/popup.js';
import { sheetScrollFade } from '../../styles/sheet-scroll.css';
import { PHONE_QUERY } from '@utils/breakpoints';
import { nameDialogsIn } from '@utils/name-dialog';
@@ -164,46 +165,17 @@ export class MenuSurface extends LitElement {
and worse when the cut lands on a row boundary, where the
sheet ends in a clean edge that reads as the end of the list.
Two layers, and the *order* is what asks the question: a
shadow pinned to the bottom of the box (attachment scroll),
and over it a cover of the sheet's own colour painted at the
end of the *content* (attachment local), which therefore
scrolls up over the shadow and hides it exactly when there is
nothing more to see. So the affordance is absent on a menu
that fits, present the moment one does not, and gone again at
the end of the list -- with no scroll listener, no
measurement, and nothing reaching into wa-dialog's shadow
root for the scroller. background-attachment is Chrome 4;
the reference device is Chrome 113.
**The curve is steep because the rows under it stay live.**
A scrim over a menu item is that item's text surface, and
this app's rule is that text clears 4.5:1 on every surface it
can sit on -- which the light ramp, whose bgElevated is
#e9ecef, is what makes non-theoretical. A row is 48px with
its label centred, so 32px of scrim that is already down to
a quarter strength at 14px reaches y-centre at about 0.06 and
spends its weight on the strip below the last legible label.
Measured on the dark ramp at x=300, flat 52,58,64 throughout
before: 50,56,62 at y=330, 33,37,40 at y=350 and 22,24,27 at
the bottom edge, and flat again at the end of the list. The
light ramp puts 9.9:1 on the last label. */
The two layers that say it live in styles/sheet-scroll.css
(#210), because the phone has a second sheet -- bottom-nav's
"More" -- which overflows for the same reason and must not
arrive at its own answer for what a fold looks like. What is
local to this sheet is the colour the cover is painted in:
the menus' elevated grey, handed over as --yj-sheet-surface
on the same box. */
wa-dialog::part(body) {
padding: 0;
overflow-y: auto;
background:
linear-gradient(
var(--yj-bg-elevated, #343a40),
var(--yj-bg-elevated, #343a40)
)
bottom / 100% 32px no-repeat local,
linear-gradient(
to top,
rgba(0, 0, 0, 0.6) 0%,
rgba(0, 0, 0, 0.25) 45%,
rgba(0, 0, 0, 0) 100%
)
bottom / 100% 32px no-repeat scroll;
--yj-sheet-surface: var(--yj-bg-elevated, #343a40);
${sheetScrollFade}
}
/* A sheet is dragged at with a thumb, so it says where its top
@@ -2201,11 +2201,25 @@ export class QueuePanel
`
: nothing}
</div>
<!-- **Every action here is named by aria-label**, like
the close button #24 added beside them (#170). A
title alone *is* a name, which is why a sweep for
empty names reports these clean and why an
assertion by role and name is green either way --
but it is the weakest one: title is the last
fallback in the accname order, so any content put
inside the button later silently outranks it, and
a phone has no hover to show it as a tooltip.
The titles stay. On a desktop they are the tooltip
for an icon-only control, which is a different job
from naming it, and aria-label does not do it. -->
<div class="header-actions">
<button
class="header-action-button"
@click=${() => void this.handleClearQueue()}
?disabled=${tracks.length === 0}
aria-label="Clear queue"
title="Clear queue"
>
<wa-icon
@@ -2216,6 +2230,7 @@ export class QueuePanel
class="header-action-button add-to-playlist-button"
@click=${this.handleAddToPlaylist}
?disabled=${tracks.length === 0}
aria-label="Add queue to playlist"
title="Add queue to playlist"
>
<wa-icon
+29 -7
View File
@@ -32,6 +32,18 @@ export interface ColumnDef {
id: string;
/** Human-readable header label. */
label: string;
/**
* Whether Settings may offer this column. Defaults to true.
*
* A definition is not the same thing as a *choice*. `titleArtist`
* is the phone's stacked column, picked by width in
* `PHONE_COLUMN_IDS`, and `tracklist.AllColumnIDs` in Go does not
* list it — so a tick in the configurator sends a column set the
* backend rejects with `unknown track-list column ID`, the tick
* reverts on the next render, and the only trace is a
* `console.error` (#197).
*/
configurable?: boolean;
/** Extracts the display value from a track. */
accessor: (track: library.Track) => string;
/** Default CSS width (used when no saved width exists). */
@@ -98,11 +110,16 @@ export const COLUMN_DEFS: Record<string, ColumnDef> = {
},
titleArtist: {
id: 'titleArtist',
// Named for what it sorts by, since that is the only place the
// label is user-visible: the phone has no column headers, and
// the page header's sort list is built from the *configured*
// columns rather than the drawn ones.
// Named for what it sorts by. That label is drawn nowhere
// today: the phone has no column headers, and the page header's
// sort list is built from the *configured* columns, which this
// one can never be — see `configurable` below.
label: 'Track Name',
// Chosen by width, never by the user, and rejected by the
// backend if it ever were. #197: Settings listed it anyway, so
// there were two rows called "Track Name" and the second one
// could not be selected.
configurable: false,
accessor: (t) => t.TrackName,
defaultWidth: '1fr',
comparator: (a, b) => compareStr(a.TrackName, b.TrackName),
@@ -266,10 +283,15 @@ export const COLUMN_DEFS: Record<string, ColumnDef> = {
};
/**
* All column IDs in default display order.
* Used by the settings UI to list available columns.
* The column IDs Settings may offer, in default display order.
*
* Not every definition is one: a column the user cannot choose has no
* row in the configurator, because a checkbox that cannot change
* anything is worse than an absent one — see `ColumnDef.configurable`.
*/
export const ALL_COLUMN_IDS: string[] = Object.keys(COLUMN_DEFS);
export const CONFIGURABLE_COLUMN_IDS: string[] = Object.keys(
COLUMN_DEFS,
).filter((id) => COLUMN_DEFS[id]?.configurable !== false);
/**
* Column IDs that are always searched regardless of visibility.
+68
View File
@@ -0,0 +1,68 @@
import { css } from 'lit';
/**
* A bottom sheet whose body scrolls says so, in one rule both sheets
* read.
*
* The app has two sheets — `menu-surface`'s context menu (#60) and
* `bottom-nav`'s "More" navigation (#71) — and both are capped at 85vh,
* because a surface covering the whole screen is a page rather than a
* sheet. So both overflow, and both used to overflow *silently*: the
* menu at 424x439 with eight items ending at y=470 (#207), the nav
* sheet at the same viewport with `scrollHeight` 412 against
* `clientHeight` 373 (#210). Where the cut lands on a row boundary the
* sheet ends in a clean edge that reads as the end of the list.
*
* The mechanism is #207's and is unchanged by being shared: two
* background layers on the scrolling box, whose *attachments* are the
* conditionality. A cover of the sheet's own colour is painted at the
* end of the *content* (`local`) over a shadow pinned to the box
* (`scroll`), so the cover scrolls up over the shadow exactly when
* there is nothing more to see. The fade is therefore absent on a sheet
* that fits, present the moment one does not, and gone again at the end
* of the list — with no scroll listener, no measurement and nothing
* reaching into another component's shadow root for the scroller.
* `background-attachment` is Chrome 4; the reference device is
* Chrome 113.
*
* Three things about it are load-bearing.
*
* **The cover takes the sheet's own colour, from a custom property.**
* The two sheets are different greys — the nav sheet paints
* `--yj-bg-surface`, because it holds the sidebar and two greys in one
* sheet is a seam across the middle of it, while the context sheet
* paints the menus' `--yj-bg-elevated`. A shared rule that hard-coded
* either would put that seam back on the other one, so the host sets
* `--yj-sheet-surface` on the same box and this reads it.
*
* **The curve is steep because the rows under it stay live.** A scrim
* over a menu item is that item's text surface, and this app's rule is
* that text clears 4.5:1 on every surface it can sit on — which the
* light ramp, whose `bgElevated` is `#e9ecef`, makes non-theoretical. A
* row is 48px with its label centred, so 32px of scrim already down to
* a quarter strength at 14px spends its weight on the strip below the
* last legible label: measured at 9.9:1 on that label on the light ramp,
* against 5.0:1 for a linear 48px draft at 0.8. The dark-ramp pixel
* table is in `.planning/NOTES.md` (2026-08-23).
*
* **The box is declared a scroller here too.** `overflow-y: auto` is
* part of the same statement rather than left to each host: a fade over
* a box that is not the scroller is a fade that never moves, and the
* component tier asserts the pair together for that reason.
*/
export const sheetScrollFade = css`
overflow-y: auto;
background:
linear-gradient(
var(--yj-sheet-surface, #343a40),
var(--yj-sheet-surface, #343a40)
)
bottom / 100% 32px no-repeat local,
linear-gradient(
to top,
rgba(0, 0, 0, 0.6) 0%,
rgba(0, 0, 0, 0.25) 45%,
rgba(0, 0, 0, 0) 100%
)
bottom / 100% 32px no-repeat scroll;
`;
@@ -243,6 +243,62 @@ describe('bottom-nav', () => {
expect(getComputedStyle(body).overscrollBehaviorY).toBe('contain');
});
it('says where the fold is, in the sheet\'s own colour', async () => {
const el = await fixture<Nav>('bottom-nav');
const drawer = shadow<HTMLElement & { open: boolean }>(el, 'wa-drawer');
if (!drawer) throw new Error('no drawer');
const shown = once(drawer, 'wa-after-show');
shadow<HTMLButtonElement>(el, '[data-testid="tab-more"]')?.click();
await shown;
const body = drawer.shadowRoot?.querySelector('[part~="body"]');
if (!body) throw new Error('no body part to scroll');
const style = getComputedStyle(body);
// #210. This list does not fit the phone — measured at 424x439,
// `scrollHeight` 412 against `clientHeight` 373 with the seed's
// eight destinations — and said nothing about it, which where the
// cut lands on a row boundary reads as the end of the list.
//
// The mechanism is #207's and is asserted the same way: the pair of
// attachments *is* the feature. A cover of the sheet's own colour
// painted at the end of the content (`local`) over a shadow pinned
// to the box (`scroll`), so the fade is absent on a sheet that
// fits, present the moment one does not, and gone again at the end.
expect(
style.backgroundAttachment,
'the cover must be local and the shadow must not',
).toBe('local, scroll');
expect(style.backgroundPosition).toBe('50% 100%, 50% 100%');
expect(style.backgroundSize).toBe('100% 32px, 100% 32px');
// And the colour is the local half of a shared rule: this sheet
// paints the sidebar's `--yj-bg-surface` (#212529) rather than the
// menus' elevated grey, or the fade draws the *other* sheet's
// colour across the bottom of this one — which is the seam a
// shared fragment would otherwise reintroduce.
expect(style.backgroundImage).toMatch(
/^linear-gradient\(rgb\(33, 37, 41\), rgb\(33, 37, 41\)\)/,
);
// And nothing paints over it. The sidebar's host carries the same
// grey, which inside the sheet is a second opaque copy of the
// surface drawn on top of these layers -- measured at 424x439 with
// the rule removed, the last 32px read a flat 52,58,64 with 39px
// still below, so the fade was painted and covered. That is
// `.context-menu-panel[data-sheet]`'s transparency, one sheet over.
const sidebar = shadow<HTMLElement>(el, 'app-sidebar');
if (!sidebar) throw new Error('no sidebar');
expect(getComputedStyle(sidebar).backgroundColor).toBe('rgba(0, 0, 0, 0)');
});
it('gives the sheet the whole width, which the sidebar does not take', async () => {
const el = await fixture<Nav>('bottom-nav');
@@ -0,0 +1,123 @@
/**
* #175: the first-run wizard's dismissal follows the library existing,
* not its own button being pressed.
*
* The wizard is a modal that blocks every pointer event, so a library
* arriving by another route — Settings, a direct call — used to leave
* it up over an app that was already set up. `LibraryAdded` is emitted
* by `AddLibrary` whoever calls it, which is what makes one
* subscription the whole fix.
*/
import { beforeEach, describe, expect, it } from 'vitest';
import { Events } from '../../src/events';
import { emit, stub } from '../support/harness';
import { fixture, shadow, shadowAll } from '../support/render';
import { wails } from '../support/wails-fake';
import '@components/first-run-wizard/first-run-wizard';
import type { FirstRunWizard } from '@components/first-run-wizard/first-run-wizard';
/** A library row, as `GetAllLibrariesWithTrackCounts` returns one. */
const aLibrary = {
id: 1,
name: 'Music',
path: '/home/logan/Music',
trackCount: 9,
};
/** Mount the wizard on a fresh install: no libraries yet. */
async function wizardOnAFreshInstall(): Promise<FirstRunWizard> {
stub('library.Library.GetAllLibrariesWithTrackCounts', []);
return fixture<FirstRunWizard>('first-run-wizard');
}
/** Whether the wizard is rendering its modal at all. */
function isShowing(el: FirstRunWizard): boolean {
return shadow(el, 'wa-dialog') !== null;
}
beforeEach(() => {
stub('library.Library.AddLibrary', aLibrary);
});
describe('first-run-wizard', () => {
it('shows on a fresh install and stays up until a library exists', async () => {
const el = await wizardOnAFreshInstall();
expect(isShowing(el)).toBe(true);
});
it('stays hidden when a library is already configured', async () => {
stub('library.Library.GetAllLibrariesWithTrackCounts', [aLibrary]);
const el = await fixture<FirstRunWizard>('first-run-wizard');
expect(isShowing(el)).toBe(false);
});
it('dismisses when a library appears by another route', async () => {
const el = await wizardOnAFreshInstall();
expect(isShowing(el)).toBe(true);
emit(Events.LibraryAdded, aLibrary);
await el.updateComplete;
expect(isShowing(el)).toBe(false);
});
it('does not raise itself when a library arrives while it is asking', async () => {
// The read is still in flight when the event lands, so its
// answer — an empty list — is stale by the time it returns.
let answer: (libraries: unknown[]) => void = () => {};
stub(
'library.Library.GetAllLibrariesWithTrackCounts',
() =>
new Promise((resolve) => {
answer = resolve;
}),
);
const el = await fixture<FirstRunWizard>('first-run-wizard');
emit(Events.LibraryAdded, aLibrary);
answer([]);
await el.updateComplete;
await new Promise((r) => setTimeout(r, 0));
await el.updateComplete;
expect(isShowing(el)).toBe(false);
});
it('still dismisses through its own Get Started button', async () => {
stub('frontendutil.FrontendUtil.HasNativeDirectoryPicker', true);
stub('frontendutil.FrontendUtil.DirectoryPicker', '/home/logan/Music');
const { resetDirectoryPickerCache } = await import(
'@utils/pick-directory'
);
resetDirectoryPickerCache();
const el = await wizardOnAFreshInstall();
const [choose, finish] = shadowAll<HTMLButtonElement>(el, '.btn');
choose?.click();
await new Promise((r) => setTimeout(r, 0));
await el.updateComplete;
finish?.click();
await new Promise((r) => setTimeout(r, 0));
await el.updateComplete;
expect(
wails.calls.filter((c) => c.path === 'library.Library.AddLibrary'),
).toHaveLength(1);
expect(isShowing(el)).toBe(false);
});
});
@@ -254,6 +254,15 @@ describe('menu-surface', () => {
// Both sit at the bottom, or the cover hides nothing.
expect(style.backgroundPosition).toBe('50% 100%, 50% 100%');
expect(style.backgroundSize).toBe('100% 32px, 100% 32px');
// The layers are shared with `bottom-nav`'s sheet since #210, and
// the colour is what each host still says for itself: this one
// paints the menus' `--yj-bg-elevated` (#343a40). A shared rule
// that hard-coded one grey would draw a seam across the other
// sheet, which is why the fragment reads a custom property.
expect(style.backgroundImage).toMatch(
/^linear-gradient\(rgb\(52, 58, 64\), rgb\(52, 58, 64\)\)/,
);
});
/**
@@ -0,0 +1,190 @@
/**
* Settings offers the columns the backend will accept, and no others.
*
* The list is built from `COLUMN_DEFS`, which is the *drawing* table:
* every definition the track list knows how to render, including
* `titleArtist` — the phone's stacked column, chosen by width in
* `PHONE_COLUMN_IDS` and never by a person. `tracklist.AllColumnIDs` in
* Go does not list that id, so the configurator offered a nineteenth
* row that could not be ticked:
*
* ```
* validate = unknown track-list column ID: "titleArtist"
* titleArtist valid = false
* ```
*
* What a user saw was **two rows both called "Track Name"** (#197), one
* of which did nothing — and a screen reader heard "Show the Track Name
* column" twice with nothing to tell them apart, which is `a11y.32`'s
* complaint inside the list that was fixed for exactly that.
*
* It is worse than an inert control, which is why the duplicate name
* was not the thing to fix. `SetTrackListColumns` assigns before it
* validates, so a rejected list stays in memory and `Save()` validates
* the whole config:
*
* ```
* later, unrelated SetThemeAccentColor = could not save config: invalid
* config: ... unknown track-list column ID: "titleArtist"
* ```
*
* — one tick and no setting saves for the rest of the session. That
* half is filed separately; this file keeps the row from being offered.
*
* The last test is the one that would have caught it when the column
* was added: the two lists are in different languages, so nothing but a
* sweep can hold them together.
*/
import { beforeEach, describe, expect, it } from 'vitest';
import '@components/config-page/config-page';
import {
COLUMN_DEFS,
CONFIGURABLE_COLUMN_IDS,
} from '@components/track-list/columns';
import { flush, stub } from '@test/support/harness';
import { fixture, shadowAll } from '@test/support/render';
/** Go's own list of column ids, as text. */
const GO_CONFIG = Object.values(
import.meta.glob<string>('../../../backend/tracklist/config.go', {
eager: true,
query: '?raw',
import: 'default',
}),
)[0];
/**
* The ids `tracklist.AllColumnIDs` actually contains.
*
* Read out of the source rather than written down here, because a
* third copy of this list is a third thing to forget — which is the
* defect, one copy earlier.
*/
function goColumnIDs(source: string): string[] {
const constants = new Map<string, string>();
const constBlock = /const \(([\s\S]*?)\n\)/.exec(source)?.[1] ?? '';
for (const [, name, id] of constBlock.matchAll(
/(\w+)\s+ColumnID\s*=\s*"([^"]+)"/g,
)) {
constants.set(name!, id!);
}
const listBlock =
/var AllColumnIDs = \[\]ColumnID\{([\s\S]*?)\n\}/.exec(source)?.[1] ?? '';
return [...listBlock.matchAll(/(\w+),/g)]
.map(([, name]) => constants.get(name!))
.filter((id): id is string => id !== undefined);
}
/**
* The column rows, and only those.
*
* Settings view-visibility list (#25) is drawn with the same two
* classes, so a bare `.column-label` sweeps 29 rows across two
* sections — and "Albums" the destination sitting beside "Album" the
* column is not the fault this file is about. The `for`/`id` prefix is
* what tells them apart.
*/
const COLUMN_ROW_LABEL = 'label.column-label[for^="column-"]';
const COLUMN_ROW_BOX = 'input.column-toggle[id^="column-"]';
/** The rows the configurator draws, by their visible name. */
async function columnRowNames(): Promise<string[]> {
const page = await fixture('config-page');
await flush();
await page.updateComplete;
// Every section renders collapsed, and a collapsed body is `hidden`.
for (const section of shadowAll<HTMLElement>(page, 'config-section')) {
section.shadowRoot
?.querySelector<HTMLButtonElement>('button[aria-expanded="false"]')
?.click();
}
await flush();
await page.updateComplete;
return shadowAll<HTMLElement>(page, COLUMN_ROW_LABEL).map(
(label) => label.textContent?.trim() ?? '',
);
}
describe('the Settings column list', () => {
beforeEach(() => {
for (const path of [
'library.Library.GetAllLibrariesWithTrackCounts',
'jobs.Service.GetJobs',
'download.Service.ListProviders',
'download.Service.ProviderKinds',
]) {
stub(path, []);
}
stub('config.Config.GetShortcuts', {});
stub('config.Config.GetDownloadPreferences', {});
stub('config.Config.GetThemeAccentColor', '#ffd43b');
stub('config.Config.GetThemeBackgroundShade', 'dark');
});
it('names each row once', async () => {
const names = await columnRowNames();
// A sweep over nothing passes.
expect(names.length, 'the page draws column rows').toBeGreaterThan(5);
const seen = new Set<string>();
const duplicated = names.filter((name) => !seen.add(name));
expect(duplicated).toEqual([]);
expect(names.filter((n) => n === 'Track Name')).toHaveLength(1);
});
it('gives each checkbox a name that identifies it', async () => {
// The visible half above is what was reported; this is the half a
// screen reader gets, and it is the one `config-page` computes
// from the same string.
const page = await fixture('config-page');
await flush();
await page.updateComplete;
const labels = shadowAll<HTMLInputElement>(page, COLUMN_ROW_BOX).map(
(box) => box.getAttribute('aria-label') ?? '',
);
expect(labels.length, 'the page draws column checkboxes').toBeGreaterThan(5);
expect(new Set(labels).size).toBe(labels.length);
});
});
describe('the column table', () => {
it('offers no column the backend would reject', async () => {
const accepted = goColumnIDs(GO_CONFIG ?? '');
// Two non-vacuity guards: a glob that stopped matching, and a
// parse that stopped finding the list it names.
expect(GO_CONFIG, 'backend/tracklist/config.go is readable').toBeTruthy();
expect(accepted.length, 'AllColumnIDs was parsed').toBeGreaterThan(10);
expect(
CONFIGURABLE_COLUMN_IDS.filter((id) => !accepted.includes(id)),
).toEqual([]);
});
it('still knows how to draw every column it offers', async () => {
// The filter must not have taken a column *out* of the drawing
// table: `configurable` says what Settings may list, not what the
// list may render.
expect(
CONFIGURABLE_COLUMN_IDS.filter((id) => COLUMN_DEFS[id] === undefined),
).toEqual([]);
expect(CONFIGURABLE_COLUMN_IDS).not.toContain('titleArtist');
expect(COLUMN_DEFS['titleArtist'], 'the phone still has its column')
.toBeTruthy();
});
});
+7 -3
View File
@@ -38,10 +38,14 @@ pre-commit:
glob: "*.go"
run: ./scripts/bindings-check.sh
# .pi/ documents make targets; a stale one sends an agent off a
# cliff with total confidence. Instant.
# The docs document make targets; a stale one sends an agent — or a
# contributor reading CONTRIBUTING.md — off a cliff with total
# confidence. The glob is the script's own scanned set, because a
# hook that does not fire on a file the check reads is the drift the
# check exists to prevent: it was `{Makefile,.pi/**/*.md}` while the
# script already read CLAUDE.md. Instant.
skill-check:
glob: "{Makefile,.pi/**/*.md}"
glob: "{Makefile,.pi/**/*.md,AGENTS.md,CLAUDE.md,README.md,CONTRIBUTING.md}"
run: ./scripts/skill-check.sh
frontend-typecheck:
+21 -4
View File
@@ -14,6 +14,11 @@
# missing: CLAUDE.md names 27 targets and nothing verified one of them,
# so the file the agents trust most was the file least checked.
#
# README.md and CONTRIBUTING.md are in it too, and the header sentence
# above is why: a person who has *not* read the Makefile goes looking in
# the contributor-facing doc, so a renamed target sends them off the
# same cliff it sends an agent off. CONTRIBUTING.md names 21 targets.
#
# **AGENTS.md is a symlink to CLAUDE.md.** This repo is worked on by
# two agent harnesses that read different files by convention — Claude
# Code reads CLAUDE.md, others read AGENTS.md — and two harnesses
@@ -43,7 +48,19 @@ if [ -e AGENTS.md ] || [ -L AGENTS.md ]; then
fi
fi
[ -d .pi ] || exit 0
# The scan is over the docs that are actually there: a checkout without
# .pi/ still has README.md and CONTRIBUTING.md to check, and gating the
# whole run on .pi/ would have made the human-facing half conditional on
# the agent-facing one. This list is used twice — once to read the
# mentions out and once to say which file a missing target came from —
# because a second list is a second thing to forget.
# `ls` exits non-zero when *any* of its arguments is missing while still
# printing the ones that are there, and under `set -e` that would sink
# the assignment rather than scanning what exists, so swallow it.
docs="$({ find .pi -name '*.md' 2>/dev/null
ls CLAUDE.md README.md CONTRIBUTING.md 2>/dev/null || true; })"
[ -n "$docs" ] || exit 0
# `make -pq` prints the database including every rule, without running
# anything. It exits non-zero when a target is out of date, and under
@@ -68,7 +85,7 @@ targets="$({ make -pqRr 2>/dev/null || true; } |
# AGENTS.md is deliberately not in this list: it is a symlink to
# CLAUDE.md, asserted above, so scanning it would report every failure
# twice under two names.
mentioned="$({ find .pi -name '*.md' 2>/dev/null; echo CLAUDE.md; } |
mentioned="$(printf '%s\n' "$docs" |
xargs awk '
FNR == 1 { fence = 0 }
/^```/ { fence = !fence; next }
@@ -93,10 +110,10 @@ for t in $mentioned; do
done
if [ -n "$missing" ]; then
echo "skill-check: the agent docs name make targets that do not exist:" >&2
echo "skill-check: the docs name make targets that do not exist:" >&2
for t in $missing; do
echo " make $t" >&2
grep -rln "make $t" .pi CLAUDE.md --include='*.md' | sed 's/^/ /' >&2
printf '%s\n' "$docs" | xargs grep -ln "make $t" | sed 's/^/ /' >&2
done
echo "Fix the docs, or restore the target." >&2
exit 1