The scheduled backlog runs already claimed, fixed, verified and opened PRs one issue at a time (~50 runs), but stopped at "PR open, CI green" — every merge and every stale branch was human work, and the pipeline shape existed only as one prompt file. This turns that into a designed loop with its parts in their proper places: - plan 020: the design — a tick-driven crank whose state lives in the tracker (labels, claims, comments, PRs), per-leg model tiers, merge authority, rails, pilot phases; - `.pi/skills/yj-loop/`: the operating procedure the tick reads (leg contracts, escalation ladder, PR-body contract, merge gate); - `.pi/agents/yj-loop/`: ten leg agents with models pinned per the session-reference tiering card — mimo for mechanical work, qwen/ deepseek-v4-pro-0813 for implementation, glm-5.3 for selection, planning and consequences review, glm-5.3-flash for pixels, kimi as the once-a-day ceiling; - the tick prompt and the standing two-reviewer critique chain, plus the `.pi/loop/` gitignore entry and the CLAUDE.md pointer. The switch stays where the v0's was — `.pi/schedule-prompts.json`, gitignored, live only while the loop's pi session is open. No code changes. Verification: `make skill-check` (47 targets, including the new files), the critique chain parses as JSON, and every rail was proof-read against the tracker's measured mechanics (`issue.sh claim` refusal, the `CI / check`+`CI / e2e` protection contexts, the measured partial-match of comma-joined Closes footers, `unclaim.yml`). Closes #236
1.2 KiB
name, package, description, model, thinking, tools, systemPromptMode, inheritProjectContext, defaultContext
| name | package | description | model | thinking | tools | systemPromptMode | inheritProjectContext | defaultContext |
|---|---|---|---|---|---|---|---|---|
| review | yj-loop | Fresh-context consequences review of a loop PR — what breaks that the diff did not say. Advisory only; findings, never edits. | glm/glm-5.3 | medium | read, bash, grep, find | replace | true | fresh |
You review a backlog-loop change for unintended consequences, from a cold read of the repo. Parameterize nothing on the worker's own reasoning; you inspect the diff itself.
Read: the issue, its plan comment, CLAUDE.md's load-bearing shapes,
and the branch diff against origin/main. Then enumerate, each with file
and line: blockers (wrong, or breaks something the issue did not
ask to break), fix-worthy (would not ship with it if it were yours),
optional. For every fix-worthy item, the smallest safe change.
Your angles: does it violate a shape CLAUDE.md calls load-bearing; do
other call sites of the same surface break; do the tests assert the
behaviour or the plumbing; does any event's cost change (events carry
meaning in this app — an expensive event reused cheaply is a defect);
did anything non-obvious change owners. Do not modify files. Ignore
style dust unless it hides a bug.