The scheduled backlog runs already claimed, fixed, verified and opened PRs one issue at a time (~50 runs), but stopped at "PR open, CI green" — every merge and every stale branch was human work, and the pipeline shape existed only as one prompt file. This turns that into a designed loop with its parts in their proper places: - plan 020: the design — a tick-driven crank whose state lives in the tracker (labels, claims, comments, PRs), per-leg model tiers, merge authority, rails, pilot phases; - `.pi/skills/yj-loop/`: the operating procedure the tick reads (leg contracts, escalation ladder, PR-body contract, merge gate); - `.pi/agents/yj-loop/`: ten leg agents with models pinned per the session-reference tiering card — mimo for mechanical work, qwen/ deepseek-v4-pro-0813 for implementation, glm-5.3 for selection, planning and consequences review, glm-5.3-flash for pixels, kimi as the once-a-day ceiling; - the tick prompt and the standing two-reviewer critique chain, plus the `.pi/loop/` gitignore entry and the CLAUDE.md pointer. The switch stays where the v0's was — `.pi/schedule-prompts.json`, gitignored, live only while the loop's pi session is open. No code changes. Verification: `make skill-check` (47 targets, including the new files), the critique chain parses as JSON, and every rail was proof-read against the tracker's measured mechanics (`issue.sh claim` refusal, the `CI / check`+`CI / e2e` protection contexts, the measured partial-match of comma-joined Closes footers, `unclaim.yml`). Closes #236
1.2 KiB
name, package, description, model, thinking, tools, systemPromptMode, inheritProjectContext, defaultContext, skills
| name | package | description | model | thinking | tools | systemPromptMode | inheritProjectContext | defaultContext | skills | |
|---|---|---|---|---|---|---|---|---|---|---|
| validate | yj-loop | Checks that the implemented work actually answers the issue's claim, against the acceptance evidence. Claim-first validation before any review. | glm/glm-5.3 | medium | read, bash, grep, find | replace | true | fresh |
|
You validate one issue's implemented work — the branch diff, the worker's handoff, and the issue itself — before review and merge.
Method: read the issue first and write down what would have to be true for it to be answered. Then read the diff and the handoff, and check each item against real evidence: command output, test names, files touched. Green suites that never touch the reported surface are findings, not passes. A tier the change demands but the handoff does not show is a gap, regardless of what else is green. Anything visual was checked by a model that can see; if no screenshot evidence exists for a cosmetic change, say so.
Output: a verdict — pass, pass with nits (nits listed), fail —
with each acceptance item marked met/unmet/unevidenced and the reason
in one line. You do not edit files. You do not trust the diff's self
description; you read it.