Files
yellowjacket/.gitea/workflows/index-artifact.yml
T
yonluandClaude Opus 5 e51cb13662
CI / check (pull_request) Canceled after 0s
CI / e2e (pull_request) Canceled after 0s
CI / check (push) Canceled after 0s
CI / e2e (push) Canceled after 0s
ci: trigger the catalog job deliberately, pin agent docs to one file
Two guardrails for the 2026-08-17 incident, and one is not about CI.

index-artifact.yml's `push` trigger was commented out that day with a
note to restore it once the rebuild completed. Restoring it is the bug.
A refresh is individually cheap, which is what made the trigger look
free; what it actually did was put an unattended job that mutates the
only copy of a ~205 GB catalog on the same trigger as an ordinary code
change, on a runner with capacity 1. The rule the file now states is the
general one -- a job that mutates state which cannot be rebuilt in ten
minutes is triggered deliberately -- so the next such job has somewhere
to look. The cron and workflow_dispatch lose nothing: indexbuild resumes
from its checkpoint either way.

Note what no branching or PR gate would have caught here. That change
was green on its branch, green on the merge and green on main; the fault
existed only against the persistent /cache database, which no fixture
reproduces. Code is gated by CI, irreplaceable state by refusing to
touch it and by docs/index-cache.md's restore.

The other half is the mismatch that started this: two harnesses reading
two files. AGENTS.md is a symlink to CLAUDE.md and skill-check asserts
the symlink rather than comparing contents, because a copy would satisfy
every other check in this repo while silently drifting -- which is the
failure being prevented. The same check now scans CLAUDE.md for make
targets, which it never did: 27 targets named in the file agents trust
most, none of them verified. Coverage goes 19 -> 46.

Scanning prose meant the line-start rule needed a fence. "Two green
branches do not / make a green merge" wrapped onto a line beginning
`make a` and duly failed on a target called `a`. Inside a fence it is
code; outside one it is a sentence that broke there, and a check that
fails on reflow gets disabled rather than fixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AfVYUVExXsx1nSWrXN8mAh
2026-08-17 15:30:38 -04:00

188 lines
7.9 KiB
YAML

name: Search index maintenance
# indexbuild decides what to do from the index's own state, so every
# trigger below runs the same command:
#
# no completed import -> build (first run, or resume a partial one)
# import older than 6mo -> rebuild (re-import from the newest dump)
# otherwise -> refresh (fold in new incremental listens)
#
# **There is deliberately no `push` trigger, and restoring one is a
# decision rather than a cleanup.** A refresh is individually cheap, so
# running it on every push to main looked free; what it actually does is
# put an unattended job that mutates the only copy of a ~205 GB catalog
# on the same trigger as an ordinary code change, on a runner with
# capacity 1.
#
# That is not hypothetical. On 2026-08-17 `fix(database): retire a table
# whose shape the schema moved past` landed on main, green — the CI
# database is deliberately in the older encoding, so the stale-shape
# repair judged its `explore_index` stale and dropped it, and this job
# fell back to a full import from the dumps. `fix(database): never
# retire the catalog the index build derives` stops that specific repair
# and cannot undo it. Every push to main then booked another `budget`
# (3h) of the one runner while ordinary CI queued behind it.
#
# So the rule this file is an instance of: **a job that mutates state
# which cannot be rebuilt in ten minutes is triggered deliberately, not
# by a push.** The weekly cron keeps the catalog current, and
# workflow_dispatch resumes or forces a build — indexbuild picks up from
# its checkpoint either way, so nothing is lost by not running on every
# merge. See docs/index-cache.md for the snapshot and the restore.
on:
schedule:
# Weekly update pass. The 6-month rebuild is triggered by the same
# command when it notices the import has aged out.
- cron: '0 4 * * 1'
workflow_dispatch:
inputs:
mode:
description: 'auto | build | refresh | rebuild'
required: false
default: 'auto'
budget:
description: 'Max build time this run'
required: false
default: '3h'
artists:
description: 'Top artists in the core artifact'
required: false
default: '50000'
# Runs share one persistent working directory, so they must not overlap.
# A push landing mid-build waits rather than corrupting the checkpoint.
#
# That directory holds the only copy of a catalog nothing can cheaply
# re-derive: see docs/index-cache.md for the snapshot it takes and the
# restore, which is minutes against the hours a rebuild costs.
concurrency:
group: search-index
cancel-in-progress: false
jobs:
maintain-index:
runs-on: ubuntu-latest
container:
# CGO is not needed: the project uses the pure-Go modernc sqlite
# driver, and neither command imports the Wails app — which is a
# claim with a test behind it now (cmd/indexbuild/deps_test.go),
# because the v3 migration quietly broke it and this job was where
# that surfaced.
image: golang:1.25
# This host path must exist on the runner and be listed verbatim in
# act_runner's container.valid_volumes. It holds explore-staging/
# (counts.bin + state.json) and yj.db — the checkpoint that makes
# resuming possible. Losing it means re-downloading ~205GB.
volumes:
- /srv/yellowjacket/index-cache:/cache
env:
YJ_HOME: /cache
CGO_ENABLED: '0'
PACKAGE_TOKEN: ${{ secrets.PACKAGE_TOKEN }}
SERVER_URL: ${{ github.server_url }}
OWNER: ${{ github.repository_owner }}
REPO: ${{ github.repository }}
SHA: ${{ github.sha }}
MODE: ${{ inputs.mode || 'auto' }}
BUDGET: ${{ inputs.budget || '3h' }}
ARTISTS: ${{ inputs.artists || '50000' }}
steps:
# Cloned by hand rather than with actions/checkout: that is a JS
# action and needs node inside the job container, which the golang
# image does not carry. Same approach as arch-package.yml.
- name: Clone repo at this commit
run: |
set -eu
git clone --quiet \
"https://x-access-token:${PACKAGE_TOKEN}@${SERVER_URL#https://}/${REPO}.git" \
/src
git -C /src checkout --quiet --detach "$SHA"
git -C /src log --oneline -1
- name: Verify the cache volume
run: |
set -eu
mkdir -p /cache
# A RAM-backed cache would defeat the point: the checkpoint has
# to outlive the job, and the import wants real disk headroom.
fstype=$(stat -f -c %T /cache || echo unknown)
echo "cache fstype: $fstype"
case "$fstype" in
tmpfs|ramfs)
echo "::error::/cache is RAM-backed; use a disk-backed host path."
exit 1 ;;
esac
df -h /cache
- name: Build tools
working-directory: /src
# The dump importer is behind the `indexbuild` tag so it is not
# linked into the app binary; cmd/indexbuild carries the same tag
# and will not build without it.
run: |
go build -tags indexbuild -o /usr/local/bin/ ./cmd/indexbuild
go build -o /usr/local/bin/ ./cmd/indexexport
- name: Maintain index
id: maintain
run: |
set +e
indexbuild -mode "$MODE" -budget "$BUDGET"
code=$?
set -e
case "$code" in
0) ;;
3) echo "::notice::Build checkpointed with work remaining — rerun to continue." ;;
*) exit "$code" ;;
esac
# Publishing only on `changed` keeps identical artifacts from
# accumulating when a refresh finds nothing new.
- name: Export core artifact
if: steps.maintain.outputs.complete == 'true' && steps.maintain.outputs.changed == 'true'
run: |
set -eu
command -v zstd >/dev/null 2>&1 || { apt-get update -qq && apt-get install -y -qq zstd; }
indexexport -o /tmp/core-index.db -artists "$ARTISTS"
zstd -19 -T0 -q -f /tmp/core-index.db -o /tmp/core-index.db.zst
sha256sum /tmp/core-index.db.zst | tee /tmp/core-index.db.zst.sha256
ls -lh /tmp/core-index.db.zst
- name: Publish to the Gitea package registry
if: steps.maintain.outputs.complete == 'true' && steps.maintain.outputs.changed == 'true'
run: |
set -eu
pkg="${SERVER_URL}/api/packages/${OWNER}/generic/yellowjacket-core-index"
# Published twice: under a dated version for history, and under
# the fixed "latest" version the client fetches. Clients cannot
# discover the newest dated version on their own — the package
# listing API requires a token, while a plain file GET does not
# — so "latest" is what makes an anonymous first run possible.
#
# A generic package rejects re-uploading a filename that already
# exists, so "latest" is deleted before being rewritten. It is
# absent on the very first publish, hence the tolerated 404.
curl --silent --show-error --user "${OWNER}:${PACKAGE_TOKEN}" \
--request DELETE "${pkg}/latest" || true
for version in "$(date -u +%Y%m%d)" latest; do
for f in core-index.db.zst core-index.db.zst.sha256; do
echo "Uploading $f -> $version"
curl --fail-with-body --user "${OWNER}:${PACKAGE_TOKEN}" \
--upload-file "/tmp/$f" "${pkg}/${version}/${f}"
done
done
- name: Summary
if: always()
run: |
echo "complete=${{ steps.maintain.outputs.complete }}"
echo "changed=${{ steps.maintain.outputs.changed }}"
if [ "${{ steps.maintain.outputs.complete }}" != "true" ]; then
echo "Build incomplete — rerun to continue from the checkpoint."
echo "Progress lives in /cache/data/explore-staging."
elif [ "${{ steps.maintain.outputs.changed }}" != "true" ]; then
echo "Nothing new to publish."
fi