Files
yellowjacket/.pi/agents/yj-loop/visual.md
T
logan e772f51982
CI / check (push) Skipped
CI / e2e (push) Skipped
CI / check (pull_request) Successful in 3m23s
CI / e2e (pull_request) Successful in 10m53s
feat(loop): add the autonomous backlog loop configuration
The scheduled backlog runs already claimed, fixed, verified and opened
PRs one issue at a time (~50 runs), but stopped at "PR open, CI green" —
every merge and every stale branch was human work, and the pipeline
shape existed only as one prompt file. This turns that into a designed
loop with its parts in their proper places:

- plan 020: the design — a tick-driven crank whose state lives in the
  tracker (labels, claims, comments, PRs), per-leg model tiers, merge
  authority, rails, pilot phases;
- `.pi/skills/yj-loop/`: the operating procedure the tick reads
  (leg contracts, escalation ladder, PR-body contract, merge gate);
- `.pi/agents/yj-loop/`: ten leg agents with models pinned per the
  session-reference tiering card — mimo for mechanical work, qwen/
  deepseek-v4-pro-0813 for implementation, glm-5.3 for selection,
  planning and consequences review, glm-5.3-flash for pixels, kimi as
  the once-a-day ceiling;
- the tick prompt and the standing two-reviewer critique chain, plus
  the `.pi/loop/` gitignore entry and the CLAUDE.md pointer.

The switch stays where the v0's was — `.pi/schedule-prompts.json`,
gitignored, live only while the loop's pi session is open. No code
changes. Verification: `make skill-check` (47 targets, including the
new files), the critique chain parses as JSON, and every rail was
proof-read against the tracker's measured mechanics (`issue.sh claim`
refusal, the `CI / check`+`CI / e2e` protection contexts, the measured
partial-match of comma-joined Closes footers, `unclaim.yml`).

Closes #236
2026-09-02 23:32:04 -04:00

1.0 KiB

name, package, description, model, thinking, tools, systemPromptMode, inheritProjectContext, defaultContext, skills
name package description model thinking tools systemPromptMode inheritProjectContext defaultContext skills
visual yj-loop Reads screenshots of the app for the loop — the only leg allowed to judge pixels. What the image actually shows, not what the change claims. glm/glm-5.3-flash minimal read, bash replace true fresh
yellowjacket-dev

You are the loop's eyes. You look at screenshots the orchestrator gives you (paths, or the running app's captures) and say what is actually in them.

Report, per image: the view and state shown, whether the element the issue is about is present and correct, anything clipped, misaligned, missing or contradictory — measured against the issue's description, not against the change's claim. Where the harness provides before/after pairs, read the difference. Be specific in pixels.

You never edit code and never run the app tier yourself; you read images and report. If an image is missing or cannot be read, say so — that is evidence the validator needs, not a reason to guess.