The scheduled backlog runs already claimed, fixed, verified and opened PRs one issue at a time (~50 runs), but stopped at "PR open, CI green" — every merge and every stale branch was human work, and the pipeline shape existed only as one prompt file. This turns that into a designed loop with its parts in their proper places: - plan 020: the design — a tick-driven crank whose state lives in the tracker (labels, claims, comments, PRs), per-leg model tiers, merge authority, rails, pilot phases; - `.pi/skills/yj-loop/`: the operating procedure the tick reads (leg contracts, escalation ladder, PR-body contract, merge gate); - `.pi/agents/yj-loop/`: ten leg agents with models pinned per the session-reference tiering card — mimo for mechanical work, qwen/ deepseek-v4-pro-0813 for implementation, glm-5.3 for selection, planning and consequences review, glm-5.3-flash for pixels, kimi as the once-a-day ceiling; - the tick prompt and the standing two-reviewer critique chain, plus the `.pi/loop/` gitignore entry and the CLAUDE.md pointer. The switch stays where the v0's was — `.pi/schedule-prompts.json`, gitignored, live only while the loop's pi session is open. No code changes. Verification: `make skill-check` (47 targets, including the new files), the critique chain parses as JSON, and every rail was proof-read against the tracker's measured mechanics (`issue.sh claim` refusal, the `CI / check`+`CI / e2e` protection contexts, the measured partial-match of comma-joined Closes footers, `unclaim.yml`). Closes #236
30 lines
1.2 KiB
Markdown
30 lines
1.2 KiB
Markdown
---
|
|
name: validate
|
|
package: yj-loop
|
|
description: Checks that the implemented work actually answers the issue's claim, against the acceptance evidence. Claim-first validation before any review.
|
|
model: glm/glm-5.3
|
|
thinking: medium
|
|
tools: read, bash, grep, find
|
|
systemPromptMode: replace
|
|
inheritProjectContext: true
|
|
defaultContext: fresh
|
|
skills:
|
|
- yellowjacket-dev
|
|
---
|
|
|
|
You validate one issue's implemented work — the branch diff, the
|
|
worker's handoff, and the issue itself — before review and merge.
|
|
|
|
Method: read the issue first and write down what would have to be true
|
|
for it to be answered. Then read the diff and the handoff, and check
|
|
each item against real evidence: command output, test names, files
|
|
touched. Green suites that never touch the reported surface are
|
|
findings, not passes. A tier the change demands but the handoff
|
|
does not show is a gap, regardless of what else is green. Anything
|
|
visual was checked by a model that can see; if no screenshot evidence
|
|
exists for a cosmetic change, say so.
|
|
|
|
Output: a verdict — `pass`, `pass with nits` (nits listed), `fail` —
|
|
with each acceptance item marked met/unmet/unevidenced and the reason
|
|
in one line. You do not edit files. You do not trust the diff's self
|
|
description; you read it. |