One autopilot day on our own repo
Snapshot: 2026-09-29, taken while the run was still in progress. Figures cite their source: the #880 log, GitHub pull requests and releases, or the gh query shown.
What we did
On 2026-09-29 the owner set a charter for a self-paced loop on the ai-dossier repository: work on automation, product quality and observability, with full autonomy to merge and release once CI is green and the orchestrator has reviewed the diff. The loop's persistent log is a single GitHub issue, #880, so a dead session can be resumed from the last cycle comment.
- Orchestrator: one Claude Code session (Opus 5.5) that picks work, reviews and merges.
- Workers: Sonnet 5.5 subagents, one per issue or small batch, each in its own git worktree, running the repo's own
full-cycle-issuedossier (plan, implement, test, review, PR). - Guardrails: workers stop at a green-CI pull request and never merge; a budget gate stops the loop at 40% weekly Claude usage; issues claimed by another session, or labeled
in-progressorblocked, are skipped; product decisions are written up as options for the owner rather than guessed.
What came out
| Figure | Value | Source |
|---|---|---|
| Cycles completed | 10 (an 11th was in progress) | #880 cycle comments |
| Pull requests merged from the loop's cycles | 25 (from Cycle 1 to Cycle 10) | PR numbers named in the #880 cycle comments |
All PRs merged to main in the same window | 29, including 4 from a parallel scheduler session; +17,076 / −1,563 lines | GitHub search: merged since 2026-09-29, counting those after the 08:28 UTC start |
| Releases published | 11 GitHub releases: CLI cli-v0.72.0 to cli-v0.80.0 (10, including 0.76.1) and vscode-v0.1.0 | Releases page, gh release list |
| Registry dossiers republished | 3: ship-issue 1.18.0, full-cycle-issue 3.17.0, fleet-cycle 1.9.0 | Cycle 9 comment in #880 |
| Weekly Claude usage | 6% at start, 20% at the end of Cycle 9 (account-wide, so other sessions count) | Usage readings in the #880 cycle comments; the usage API returned HTTP 503 at Cycle 10 close |
Shipped items, all adopter-facing or reliability work:
- #901: standalone no-Node CLI binaries (Node SEA) with
install.sh, released for 5 platforms; verified end to end on a host with no Node on PATH. - #914: VS Code extension MVP (inline diagnostics, completion, dry-run preview), released as
vscode-v0.1.0. - #909: deep static dry-run preview with a 0-100 risk score, also exposed through the MCP server. Every surface says it is a preview, not a safety guarantee.
- #903: actionable checksum and signature failure messages, a confirmation prompt for high-risk dossiers, key export and revoke.
- #916: only OWNER, MEMBER and COLLABORATOR authors are trusted for
plan:v1artifacts posted as issue comments (a partial fix; see the open gaps below). - Scheduler fixes: #886, #898, #917, #918, #923; and a publish safety net, #881, because GitHub most likely dropped a push event and skipped a release (a hypothesis in the log, supported by the evidence, not proven).
How quality was controlled
Two independent review layers ran on the workers' pull requests: a review agent spawned by each worker (one Cycle 2 worker skipped this phase), then the orchestrator reading the diff, CI results and acceptance criteria before merging. The orchestrator's own small PRs (#906, #910) had no worker review. The orchestrator also held merges. Defects caught before they shipped, as recorded in #880:
- #917 (merge-mode detection): the worker's review agent reported 3 HIGH and 7 MEDIUM findings, including that the fix did not work on this very repository. All were addressed with tests before merge (3 further findings were withdrawn as false positives).
- #916 (author trust): the worker's adversarial review said "merge", but its HIGH finding described the core attack, and other readers still trusted any commenter. The orchestrator held the merge until those paths, a duplicated trust set and a bogus association value were fixed.
- #922 (focused test gate): review found that a renamed file hid its source workspace, so a red change could pass the gate. Fixed by diffing without rename detection.
- #914 (VS Code extension): the worker's review agent found 4 real bugs (an off-by-one line in YAML errors, extension tests that never ran in CI, completion breaking on wide JSON indents, stale diagnostics after closing a file).
- #909 (dry-run): the worker's final commit never landed because a commit hook failed with its output suppressed, yet the worker reported it done. The orchestrator noticed in review and shipped it as #911. Process change: the orchestrator now checks
git statusand unpushed commits in the worker's worktree before merging.
What failed, and what we did about it
- Flaky tests were partly a product bug. A runstate flake in #894 was a run-id collision (2 random bytes, about 1 in 65,536); fixed and covered by a deterministic test. Honest caveat from the log: neither flake reproduced locally in 90 and 60 attempts, so those fixes rest on code reading.
- Release automation needed two orchestrator fixes and is still not clean. The binaries build fired before npm's listing had caught up, and two runs raced to create the same release (HTTP 403). Fixes: #906 and #910. The 403s then recurred in two serialized runs, so "race" did not explain them; the log had mis-diagnosed it. Root cause is open in #937.
- Workers did not always follow their own procedure. One Cycle 2 worker skipped the review phase, and after a merge one worker kept editing (#922 then #925).
- Review reports were not delivered. Review agents in Cycles 9 and 10, and all four sub-reviewers of the later audit, finished without their report reaching the requester. This was harness behavior, not worker behavior, and prompt changes could not fix it; the orchestrator now owns relaying reports.
- Reviewers trusted a stale snapshot. Reviewers read
examples/, which lags the registry, and raised 3 false findings. Reviewers are now told to read the registry orpackages/. - The most common failure was "done but never committed". Besides #909 (above), another session had 16 files stranded in a worktree. The stated missing invariant: commit and push before any long wait.
- Reviewed and merged does not mean finished. A later independent 6-hour audit of the day's riskiest PRs found HIGH-severity gaps in two shipped changes; fixes were in progress at the snapshot: #932 (the scheduler still acts on runstate milestones posted by any commenter, so the author-trust work in #916 covers only part of the surface) and #933 (the dry-run analyzer from #909 under-reports some dangerous dossiers as low risk)
- Tooling friction. Harness worktree isolation fails on this repo's nested
main/.gitlayout, so workers create worktrees by hand; the permission classifier blocked unreviewed subagent merges until a scoped rule was added; a PR's CI did not start for about 25 minutes because of a merge conflict.
What stayed with humans
Product and spending decisions (website hosting, telemetry, JetBrains plugin, monetization SDK, Marketplace publishing) were posted as options with a recommendation and resolved by the owner. Accounts and secrets for signing and the VS Code Marketplace remain owner-only.
Limits of this evidence
- One repository, one day, one team, and the team is the tool's author. It is an existence proof, not a benchmark.
- We did not measure how long a human would have taken, so we make no time-saved claim.
- Lines changed and PR counts measure activity, not value. Several PRs were fixes to the automation itself.
- Some fixes could not be reproduced against a failing case, as noted above.
Source log: issue #880. Related reports in docs/reports.