Agent Traps
Grep-first symptom → trap → fix index. plan-issue reads this file in full and greps it
for terms from the issue title and affected paths before planning (Step 4.5); report-issue
appends a row here when a run required a CI fix or surfaced a trap a future agent should
search for first.
Repo and CI traps
| Symptom (grep this) | What went wrong | Fix | PR |
|---|---|---|---|
version-bump CI fails after changing packages/sched/src/ with package version must be greater than the base branch version | A source change in a publishable workspace requires its own package version bump. Tests and builds can pass locally while the release gate correctly rejects the PR because packages/sched/package.json still carries the version already on the base branch. | Before opening a PR that changes publishable source, bump that workspace (npm version patch --workspace=@ai-dossier/sched --no-git-tag-version when appropriate), regenerate package-lock.json, then run node scripts/check-version-bumps.mjs --base origin/main --head HEAD. | PR #869 (#866) |
check-version-bumps.mjs fails with a workspace dependency pin changed but the version is unchanged right after bumping ONLY the depended-on package’s version | Bumping @ai-dossier/worktree-pool and widening cli’s and packages/sched’s own @ai-dossier/worktree-pool dependency range to match is not enough on its own: npm install then resolves cli/sched’s own package.json version as unchanged, which is exactly what scripts/check-version-bumps.mjs (#692) exists to catch, and separately — before any CI run — npm install/npm ci shadow-copies the OLD published version into cli/node_modules/@ai-dossier/worktree-pool (and packages/sched/node_modules/...) instead of symlinking the local workspace package, because a 0.x caret (^0.6.x) does not satisfy 0.7.0 under semver. Both failures are the SAME root cause (a workspace-sibling repin is release-relevant on its own) surfacing at two different times: the shadow-copy at npm install time, the CI gate at PR time (#760) | Whenever a publishable package’s version bumps AND a sibling package’s dependency range on it widens to track that bump, also bump the sibling’s own package.json version (npm version patch --no-git-tag-version is usually right) and regenerate the lockfile — then verify with node scripts/check-version-bumps.mjs --base origin/main --head HEAD and confirm readlink -f node_modules/@ai-dossier/<sibling> resolves to the local packages/<sibling> path, not a .../node_modules/@ai-dossier/<sibling> shadow copy nested inside a consumer package | PR #761 (#760) |
harness produced no envelope with task-failed (exit 0) → batch suite-unreadable, while ~/.dossier/caps.jsonl shows outcome=ok exit_code=0 | cap run re-emitted a large captured stdout and called process.exit() before the pipe drained — libuv’s queued write (tail + envelope) was discarded (20 MB in, 146 KB out, exit 0). Consumers also read only stdout’s last line, so any output landing after the envelope had the same effect. The same trap bites any CLI command that writes a lot to a pipe and then calls process.exit() | Drain stdout/stderr before process.exit() (drain() in cli/src/commands/cap.ts); consumers read the envelope via spawnCapRun in cli/src/cap-envelope.ts (envelope file → marked stdout line, exit-code agreement required). An ai-dossier older than 0.59.0 on PATH still has the lossy stdout-only channel | #811 |
is the first bad commit | Newer git prints <sha> is the first bad commit on stderr, not stdout — bisect result-parsing that read stdout found nothing, so every real bisect returned error in CI while passing on older git locally | Read the result from refs/bisect/bad (git leaves it pointing at the first bad commit) instead of parsing bisect’s stdout | PR #498 |
no crontab for | set -e combined with crontab -l on a host with no existing crontab (which exits non-zero and prints no crontab for <user>) let the fleet-reset script fall through with empty captured content, so it installed an empty crontab instead of preserving the existing jobs | Capture crontab -l with || true (or 2>/dev/null || echo "") before appending new jobs, so a missing crontab reads as empty input, not a fatal error | reset-fleet incident |
unknown command | A stale nvm copy of ai-dossier — a repo-local node_modules/.bin/ai-dossier or stray ~/node_modules install — shadows the global install with an older build; running a newly documented subcommand against the stale shadow reports unknown command instead of running it Inside the repo or any worktree, PATH’s ./node_modules/.bin makes the workspace-linked cli win — its dist/ is whatever the last local build (or a stale pool worktree) left, and it reports the SAME --version as the global (read from package.json, not dist/), so a version check looks clean (error: unknown command 'runstate' on #804) | type -a ai-dossier shows which copy wins. Call the newer binary by absolute path ($(npm prefix -g)/bin/ai-dossier), make build-all in that worktree before relying on the local link, or run scripts/fleet-cli-audit.sh --fix to neutralize shadow copies (renames, never deletes) | fleet-cli-audit rationale; PR #806 (#804) |
resume_from=setup | runstate verify’s computeResume/setupOk() required the recorded worktree= path to exist on THIS machine, so a cross-machine resume (redispatch landing on a different host) silently re-ran setup and plan instead of resuming at implement, even though the branch and head= were present on origin | setupOk() now accepts remote-first evidence — branch exists on origin and head= is an ancestor of remote HEAD; local worktree presence is informational only (local_worktree=present|absent) and never gates resume_from | PR #516 (#499) |
make: unrecognized option '--reporter=json' | The batch aggregate-suite runner shelled npm test -- --reporter=json. In a repo whose root test script delegates to make ("test": "make test"), npm forwarded the flag to make, which aborted. The runner then returned {ok:false, failing:[]}, beginAttribution found nothing to attribute, and a batch whose members were all green was dissolved as unattributable-suite-failure — deterministically, every time | Fixed (#562). cli/src/batch-suite-runner.ts resolves in order — an active cap run test.full → dispatch.suite_command → a repo-detected direct invocation that never forwards extra flags through an unrecognized script. An active test.full.timeout_ms is the cap command’s budget only; its timeout is terminal (no guessed fallback), while shadow/unavailable entries and tiers 2/3 use the configured runner timeout. Never assume npm test -- forwards to a test runner | PR #556 (#526), fix #562 |
unattributable-suite-failure | Emitted when the aggregate suite is non-green but zero failing test names parsed. It dissolves the whole batch and requeues every member as an independent full-cycle run with preserved=none — so completed member work is discarded and paid for twice. An unparseable or non-runnable suite was indistinguishable here from a genuinely empty one | Fixed (#562) for the primary validating gate (runValidate). SuiteResult.readable now distinguishes “parsed a report naming zero failures” from “never got a parseable report at all” — the latter (suite-unreadable, after the fallback-runner retry when the primary tier was cap/config) blocks the batch (blockBatch, packages/sched/src/recovery.ts) instead of dissolving it: no requeue, no revert, worktree kept. Blocking has no CLI resume verb yet — sched abandon --batch is today’s only real exit. A genuinely red suite with a parseable failing-test list still dissolves exactly as before when unattributable — but note that in THIS repo (whose resolved command is make test, never a vitest JSON reporter) that path is effectively unreachable: every red aggregate suite here comes back unreadable and blocks rather than attributes; set dispatch.suite_command to a JSON-reporting invocation to keep attribution available. The post-rebase re-validate (handlePrConflict) does not yet consult readable either — a separate, narrower recovery rail, tracked as a follow-up, not fixed by #562 | PR #556 (#526), fix #562 @ai-dossier/sched >= 0.48.0 (#810): with validated members present this dissolve is refused — the batch blocks dissolve-refused:unattributable-suite-failure and the validated members stay landed. |
a sched status profile/tier shows a new model (mid=claude/sonnet) but a dispatch made by the running sched start used the OLD one (run log shows the previous model) | The engine loaded config.json once at startup and passed that object to every tick, while sched status re-reads the file on each call — so status reflected the edit and the engine never did (imboard, 2026-09-29: opus @ low -> sonnet @ medium, next member still spawned opus) | Fixed (#883, sched >= 0.63.0): the engine stats config.json + the user config each tick and reloads on change (createConfigReloader), journaling config-reloaded with the dispatch diff; an INVALID edit keeps the last good config and journals config-reload-failed (once per edit) rather than reverting to defaults. Startup-only (not reloaded): the tick interval (--interval/reconcile_interval_ms), --auto-upgrade/auto_upgrade, and the dispatch.tiers auto-detect warning. A broken USER config is reported (config-reload-failed, naming the file) and its profiles ignored, without blocking a valid project-config reload. Same-size edits within one mtime tick on a coarse-mtime filesystem can be missed (fingerprint is mtime+size). An engine older than 0.63.0 still needs a restart | issue #883 |
env-cold | A slot-cycle member hands back immediately with reason=env-cold because the batch worktree has no node_modules anywhere. claimAndSetup fetches, branches, pushes and runs git worktree add, but never runs the warmup that setup-issue-workflow’s cold path makes mandatory. Two such hand-backs exceed DISSOLVE_EVICTION_FRACTION (1/3) on a 3-member batch, so the batch dies in ~4 minutes having done no work | Fixed (#561). runBatchSetup now tries a pool claim first (already warm, mirrors teardown.ts’s poolReturn); on the cold git worktree add path it warms the worktree itself before returning — worktree.prepare via cap run when the manifest declares it, else package-manager-detected install/build (@ai-dossier/worktree-pool’s resolveWarmCommands), guarded on package.json actually being present. Journaled as batch-warmup-done/batch-warmup-failed with duration. The stopgap this superseded: hardlinking an already-warm tree in (cp -al <warm>/node_modules <batch-worktree>/node_modules) took members from 1.1-minute failures to 3/3 clean completions, proving the diagnosis | PR #556 (#526), fix #561 |
Shell cwd was reset to | On a long headless run, the shell’s working directory can be reset to the repo root between tool calls. Subsequent relative-path edits then silently target the main worktree instead of the run’s worktree — the edit either fails an anchor assertion or, worse, lands in main. It looks identical to “my edit did not apply” | Prefix every mutating command with an explicit cd <worktree> || exit 1 rather than relying on persisted cwd, and re-check pwd plus git rev-parse --abbrev-ref HEAD after any long wait | PR #556 (#526) |
events.jsonl full of byte-identical lines / Telegram flooded by one condition / a per-condition metric ~20x too high | A journal(...) / journalEvent(...) call inside a per-tick reconcile loop that continues on a DURABLE condition (unreachable ground truth, a blocked PR, a stale milestone, a merge GitHub has not reflected yet) re-emits every reconcile interval for as long as the condition lasts — one line per unit per ~2 min, burying the events an operator is actually watching for. The engine journals the evaluation instead of the transition. Fixed piecemeal three times (#610 stale-milestone-ignored, #630 pr-watch-failed, #632 the enumerated sweep of nine more), and #632’s sweep STILL missed a tenth site (runTeardownFor, #636) — the enumeration is the hard part, not the fix | Any new journal call inside a loop that does not terminate the unit needs a persisted onset marker (*_since + *_ticks on QueueEntry/BatchEntry/SlotEntry) and the shared JOURNAL_DEDUP_REANNOUNCE_TICKS re-announce window — plus a RESET at every fresh-attempt site (requeue, orphan recovery, a batch leaving awaiting-merge), or the streak silences the next run’s first occurrence. Enumerate with grep -n "journal(ctx|journalEvent(deps" packages/sched/src/{engine,batch-dispatch}.ts — BOTH files, per the sched stats row below. Emit since, not just ticks_persisted: the tick count maps to no fixed wall-clock (~20 min per-reconcile vs ~50 min on the PR-poll-gated sites, both tunable) | #610, #630, #632 (PR for #633); runTeardownFor fixed in #636. Entry-rail sites now route through the single journalConditionIfDue(ctx, state, unit, event, conditionKey, extra) / clearCondition in engine.ts (#638); pass a STABLE per-site condition key (#637), never the interpolated detail |
sched stats empty for batch members / batch-dispatch.ts | packages/sched/src/batch-dispatch.ts is a self-contained dispatcher (RFC-0001 §C.4) that never routes through engine.ts’s per-unit spawn/record functions (spawnUnit, recordDispatchRunLog) — it calls deps.spawnDeps.spawn() directly for members/tail/report/fix. Any engine-side behavior added to the per-issue dispatch path (telemetry, slot bookkeeping like spawned_at) does NOT automatically apply to batches; spawnMember’s own patch never set SlotEntry.spawned_at either, a second, independent gap found while fixing the first (a guard keyed on spawned_at !== null silently no-op’d forever) | Grep batch-dispatch.ts explicitly whenever changing anything in engine.ts’s per-unit dispatch/record/slot-patch logic — the two are duplicate, independently-maintained state machines by design (comment at the top of batch-dispatch.ts), not one dispatcher with a batch branch | #564 |
incremental-gate-failed:test.focused (batch member) / No projects matched the filters (non-batch issue-*.log only — it does NOT appear in batch gate logs; see the tee: /dev/stderr row below for why) | A batch member that completed its slot-cycle cleanly is evicted anyway. The per-member incremental gate (packages/sched/src/batch-dispatch.ts ~L1399) evicts on any cap run task-failed, and CapOutcome is derived from the exit code alone — so a suite that could not run is indistinguishable from one that ran and failed. In imboard, scripts/cap-test-focused.sh prints “Treating as inconclusive, not a pass” and exits 1 because pnpm’s --filter "...[<ref>]" git-diff selector matches zero projects inside a linked worktree — which is exactly what every batch worktree is. Every imboard batch member is therefore evicted, deterministically: #3631 (attempt 2) and #2687 (attempt 3), 2 for 2 | Partly fixed (#585). An inconclusive / automation-broken cap outcome now blocks the batch instead of evicting the member, and the gate captures its output into the unit-failed journal entry — that capture is what made the row below root-causeable in minutes. What #585 does NOT cover is a capability reporting a definite failure it did not earn — see the tee: /dev/stderr: No such device or address row below (#594 / imboard-monorepo#3996). The channel #585 added is gate-inconclusive:<capabilityId> -> blockBatch (packages/sched/src/batch-dispatch.ts), rechecked with sched resume --batch <id>. The script line it quotes (“Treating as inconclusive, not a pass”) is what cap-test-focused.sh would print if its output buffer survived — under sched dispatch it never gets there, so do not expect to grep it. This is #562’s SuiteResult.readable distinction applied to the gate #562 did not cover — rows above describe the aggregate half. Do NOT read a non-zero cap run exit as “the member’s work is red” | PR #585 (#583, closed); sched side now covered by #594 (hasEarnedFailureEvidence, packages/sched/src/attribution.ts — an unevidenced task-failed takes the same gate-inconclusive:<cap> block path); imboard-monorepo#3982 (closed); imboard-monorepo#3996 (script root cause) still open |
git-unavailable (batch member, at plan) | A slot-cycle member blocks at Step 1 because ai-dossier plan validate runs its predicted-file check as git show HEAD:<path>, which exits 128 for a path that does not exist. An issue whose entire deliverable is a NEW file therefore has a correct plan artifact that can never validate — the member is evicted having implemented nothing (imboard #1026, attempt 3) | Fixed (#579), PR #581 (c4a4740). plan validate no longer reports exit 128 (path absent at HEAD) as a git-probe failure — a predicted path that does not yet exist is a legitimate new-file prediction, not git-unavailable. Before assuming this is your bug, check the installed @ai-dossier/cli carries #581 | PR #581 (#579) |
resume_from=report on an issue you were just dispatched to work | runstate verify maps a report/done milestone + OPEN issue to resume_from=report unconditionally (the report resolver in PHASE_RESUMERS, cli/src/runstate.ts) — with no check that the milestone belongs to this dispatch. A deliberately re-enqueued issue (pilot re-run, abandon+enqueue, an issue left open because its ACs were not met) therefore tells the fresh agent to skip straight to report, post a second report for work it never did, and exit reporting success. #576 fixed the engine rail (isVerifiedComplete refuses a milestone older than dispatched_at); the gate rail disagreed with it until #582 | Fixed (#582). runstate verify now accepts --dispatched-at <iso>; a report/done milestone whose at predates it (and the issue is still OPEN) resumes resume_from=none with a freshly minted run_id and prior_run=<old> (note=stale-report-trail) instead of report. gate-issue Step 1.5 must pass --dispatched-at for this to fire — without it, verify still defaults to resume_from=report but flags the ambiguity in note rather than asserting it silently | PR #586 (#582) |
agent-exited-unverified / unverified-exit-at-strongest-tier | A dispatched agent ends its turn while a watcher is still armed — “the ci-parity Monitor is still armed and will notify when the run finishes. Nothing more to do until then” — instead of blocking on it. The engine sees a turn that ended with no verified completion, fences the run at gen=0 and redispatches one tier stronger; a second unverified exit at the strongest tier is terminal (unverified-exit-at-strongest-tier). It cost #3820 and #542 in batch-pilot attempt 2 and BOTH #2687 full-cycle generations in attempt 3 — $20.06 of spend that shipped nothing. A general dispatch failure mode, not a batch defect | Mitigated (#591). resolveDispatch now appends --disallowedTools Monitor to every claude-family dispatch argv by default (packages/sched/src/dispatch.ts) — the tool cannot be armed at all, so this is no longer purely a behavioral instruction; opt out per-project with dispatch.disallowed_tools: []. The prompt instruction (NO_BACKGROUND_EXIT_INSTRUCTION) still carries the full load for a non-claude tier (opencode, no such flag) and for agents dispatched outside sched. When the trap recurs anyway, the verify-incomplete/unit-failed journal entry now carries last_tool (parseLastToolUse, @ai-dossier/core) — check that field before opening the transcript. Still never end a turn “waiting” on an armed watcher or background job generally — this fix covers one specific tool on one specific CLI family. Attempt 4 (2026-09-03) shows it is narrowed, not closed (#596): zero occurrences across 7 batch members, but it fired on ALL THREE long-running (>30 min) full-cycle units in the same run (issue:3985 06:40:16.780Z, issue:826 06:50:18.477Z, issue:340 07:06:22.641Z) with --disallowedTools Monitor present on every dispatch — and #3985 then died unverified-exit-at-strongest-tier. Note what that did NOT mean: PR imboard-monorepo#3999 merged anyway, at 2026-09-03T07:19:58Z, and #3985 is closed — a unit failing at the strongest tier is a statement about the agent’s exit, not about whether the work shipped. Check gh pr view before writing off a failed unit’s output. The correlation is unit DURATION, not the tool. Read last_tool on the verify-incomplete entry first. Partly closed engine-side (@ai-dossier/sched >= 0.23.0, #596/#620). The manual “check gh pr view before writing off a failed unit” step above is now automatic for exactly the #3985/#3999 shape: before failing a unit terminally for an unverified exit, the engine asks gh pr list --head <branch> --state open and, on a confirmed PR the fleet itself opened, parks the unit for the watcher instead — journaled pr-parked with detail=unverified-exit-recovered-open-pr; grep that detail to find adopted units. It fails closed, and the terminal unit-failed now carries `pr_check=none | unreachable |
api_error_status / terminal_reason":"api_error" / monthly spend limit / a tail-agent-exited-unverified or unverified-exit that REPEATS on the same unit every few minutes | A provider-side 429 spend/rate wall kills the dispatch before the agent runs a single turn. The result event carries api_error_status: 429, terminal_reason: "api_error", modelUsage: {} and num_turns: 1 — but to a reconciler that only asks “did the process exit without a verified milestone”, it is byte-for-byte the same decision as an agent that ran for 40 minutes and ended its turn early. The batch tail respawned 9 times in 33 minutes against the wall before an operator killed the tick; runBatchTick never read paused, so #505’s dispatch-health pause was cosmetic for exactly this case, and the per-issue path would have burned the unit’s escalation ladder on a wall that was not the issue’s fault | Fixed (#629, @ai-dossier/core >= 1.7.0 / @ai-dossier/sched >= 0.24.0). parseDispatchApiError (@ai-dossier/core) classifies on api_error_status/terminal_reason ONLY — never the free-text result message, which is not a contract (the #609 usage/modelUsage lesson generalized); modelUsage: {} + is_error are corroboration, never the signal. dispatch-health.ts is shared by engine.ts AND all four batch reconcilers (tail/member/fix/report) precisely because of the row above — it is its own module to avoid the engine.ts ↔ batch-dispatch.ts import cycle. A confirmed error journals dispatch-failure with the provider’s status/reason/message/reset_at, feeds consecutive_dispatch_api_errors (same-unit repeats COUNT, unlike suspect-dispatch’s cross-unit rule), and at DISPATCH_UNHEALTHY_THRESHOLD reuses #505’s setPaused. The redispatch does NOT escalate — same tier, recoveries unchanged — and once paused, BOTH rails hold instead of respawning (runBatchTick’s wedges, enterRecovery/reconcileRecovering’s unescalated redispatch) rather than looping until an operator intervenes. Batch logs are fenced to log_offset_at_spawn so a later genuine crash (no result event) is never misclassified against a stale 429 from an earlier attempt to the same append-mode file. Grep dispatch-failure in events.jsonl first; sched resume clears the pause and the streak | #629 |
binary file matches from grep/rg on a .ts source, or grep silently returning NOTHING for a string you can see with sed | A literal NUL byte (0x00) written into a source file — here as a composite-map-key separator (`${model}\x00${class}`). It renders as a plain SPACE in git diff and in most editors, so review reads it as a space separator and every agent reviewing the diff misses it. file <path> reports data, GNU grep treats the whole file as binary and skips it, and ripgrep (which the agent Grep tool is built on) reports binary file matches with no lines — so the module becomes invisible to the grep-first workflow this very file depends on. tsc and biome both pass, so nothing in CI catches it. The author was an agent: an Edit tool call that intended a space wrote NUL. | Write the byte as an escape and name it: const SEP = '\u0000' (a NUL separator is the RIGHT choice — no canonical model id or enum value can contain one). To detect: file <path> should say text, and python3 -c "print(open(p,'rb').read().find(b'\x00'))" should print -1. If a grep over a file you just edited returns nothing for a string you can see, suspect this before doubting the grep. | PR #587 (#528) |
Vercel check red with Deployment has failed — run this Vercel CLI command / Resource provisioning failed + BUILD_FAILED and 0 build-log events | The dossier-registry preview deploy fails for reasons that have nothing to do with the diff — Neon’s preview-branch quota is exhausted by fleet merge throughput (#567), so provisioning fails after the build. It presents as a red required-looking check on a docs-only PR. ship-issue then spends a ci_fix_attempts pushing an empty chore: re-trigger preview deployment commit, which fails identically: a quota does not clear on retry, and retrying is what run r-526-eba7 did (ci_fix_attempts=1 for zero benefit). Production deploys succeed in the same window, which makes it look diff-specific when it is not | main in this repo has no branch protection and no required contexts (gh api repos/imboard-ai/ai-dossier/branches/main/protection → 404), so a red Vercel preview never blocks a merge. Confirm the ONLY red check is Vercel/preview (gh pr checks <n> — GitHub Actions lint/test/version-bump green), note it as ci_note=vercel-preview-...,not-a-required-check on the ship milestone, and merge. Do NOT burn a CI fix attempt on a re-trigger; if a working preview is actually needed, delete stale Neon preview branches first | PR #584 (#526), infra #567 |
test (20)/test (22) red at the Security audit step (npm audit --omit=dev --audit-level=high), never reaching Build/Run tests with coverage | A freshly-disclosed high/moderate advisory on a transitive dependency already in package-lock.json (e.g. fast-uri, qs) fails the CI job’s own npm audit gate before the PR’s actual build/tests ever run — on a PR that touches no lockfile at all. git diff origin/main...HEAD -- package-lock.json is empty and npm audit reproduces identically against origin/main’s own unchanged lockfile, proving it is not diff-specific; gh run rerun <run-id> --failed fails again identically (deterministic, not flaky — a disclosed CVE does not clear on retry) | Confirm the lockfile is untouched (git diff origin/main...HEAD -- package-lock.json empty) and reproduces on origin/main’s lockfile before treating it as a non-blocker. main has no branch protection (see the Vercel row above), so mergeStateStatus=UNSTABLE/mergeable=MERGEABLE still permits merge with lint/version-bump/test-examples green — note it (ci_note=npm-audit-preexisting-advisory,not-a-required-check) on the ship milestone and merge. Fixing the advisory (npm audit fix, a lockfile bump) is a separate PR, not this one, unless the PR being shipped already touches the lockfile | PR #593 (#591) |
tee: /dev/stderr: No such device or address in a cap run / gate log, followed by exited 1 — real test failure and a log with zero test output | scripts/cap-test-focused.sh’s run_pnpm_capture pipes through tee /dev/stderr. The gate runs the script through ai-dossier cap run, whose spawnSync captures the child’s output — and Node/libuv backs a captured stdio stream with a socketpair, so /dev/stderr (/proc/self/fd/2) names a socket, and reopening a socket that way returns ENXIO. A plain file redirect (2>somefile) and an ordinary pipe both reopen fine, so cmd 2>somefile will NOT reproduce this — check with node -e 'require("child_process").spawnSync("bash",["-c","echo hi | tee /dev/stderr"])' instead. With set -uo pipefail this (a) fabricates a non-zero exit from a pnpm run that exited 0 — pipefail returns the rightmost non-zero status of ANY element, not pnpm’s, contradicting the code comment that claims otherwise — and (b) leaves $OUT empty, so the grep -q "No projects matched the filters" guard that #3982 added never matches and its retry_with_name_filter never runs. Three members in three different batches produced 765-byte gate logs differing only in the batch id inside one path: the gate is a constant function, evicting every member regardless of its diff (5 members for 5 across attempts 2-4) | Fixed sched-side (#594) — the gate now requires evidence the capability EARNED its task-failed before evicting: failing-test output for a test.* capability, compiler errors for the others (hasEarnedFailureEvidence, packages/sched/src/attribution.ts). An empty or framing-only capture takes #585’s block-the-batch path (gate-inconclusive:<cap>) instead, on the live gate and on the sched resume --batch recheck alike; the journal detail names which branch fired. Grep for gate-inconclusive: with a detail reading reported task-failed with no failing-test evidence. Note the bar is a MARKER match, not proof: a wrapper that prints its own prose about a failure is rejected (... exited 1 — real test failure. no longer counts — an unanchored /fail/i accepted it, and the first cut of #594 shipped exactly that), but a script that fakes a FAIL line would still pass. Open — imboard-monorepo#3996 (script root cause: the tee: /dev/stderr defect itself). Diagnostic that settles it in one command: run the gate’s own pnpm --filter "...[<merge-base>]" run test by hand in the batch worktree with a normal stderr — exit 0 plus No projects matched the filters proves the suite never ran. Never read a non-zero cap run exit as “the member’s work is red” without test output in the log | #594 (fixed, @ai-dossier/sched >= 0.22.0), imboard-monorepo#3996 (open) |
batch-dissolved carrying eviction-threshold strategy=full N=<n> evictions=<m>, and you are about to conclude the batch died below its threshold | Read this before “fixing” the dissolve rule — the obvious reading is wrong. evictions=<m> in that line is ALREADY de-duplicated by distinct member (evictedCount: evictedMemberIds(batch).size), and so is the trigger itself (checkDissolveTrigger → evictedMemberIds(batch).size > threshold, packages/sched/src/recovery.ts; evictions.length appears nowhere in packages/sched/src/). De-dup landed in #572 (cfdc2bf). A requeued= entry without a matching eviction record does not BY ITSELF prove a lost eviction — strategy=full requeues every unshipped member from batch.members, evicted or not — but do not stop there: settle it from the JOURNAL. In this same batch #340 was spawned as member 3/4 and appears in only two lines of the entire journal (its own spawned, and requeued=), while the unit-failed/member-advanced pair that ended it is credited to #826, which had already failed a minute earlier. That is a MIS-ATTRIBUTED eviction, and #595’s de-dup does not fix it — appendEvictions drops the second record rather than re-pointing it, so the member’s real failure reason stays unrecoverable. Fixed in #613 (@ai-dossier/sched >= 0.22.2) — resolving a member is now a one-shot claim on executing_member, so a duplicate/re-entrant resolve no longer advances past the next member, and a lost claim journals member-advance-skipped instead of returning silently; a pre-0.22.2 state.json/events.jsonl still carries the old mis-attribution and must be read with that in mind. b-20260903-01 (N=4 evictions=3 threshold=2 requeued=47,826,340,1512) was read for three days as “2 distinct members, dissolved below threshold”; it was actually 3 distinct members (#47, #826, #1512) against threshold 2 — a CORRECT dissolve | The residual defect was narrower, and #595 fixed exactly it — NOT the dissolve rule, which was already correct. The stored evictions array retained duplicate entries (#826 twice, ~2 min apart), inflating any eviction-rate metric computed from raw event counts (RFC-0001 §E.5 reads one) even though it never moved the dissolve trigger. appendEvictions (packages/sched/src/state.ts) is now the single evictions[] append and is a no-op for a member already recorded — as is the repeat requeue that used to accompany it — journaling eviction-duplicate instead; buildStatusReport also de-dups on read, so a state.json written before #595 shows one row per member in sched status and --json. The stored array in a legacy state is left as written. Inspect the raw array before drawing any conclusion — jq '.batches[] | select(.id=="<batch-id>") | .evictions | group_by(.issue) | map({issue: .[0].issue, n: length})' ~/.dossier/sched/<project>/state.json (.batches is an ARRAY; .batches["<id>"] errors) — and check the deployed @ai-dossier/sched version carries #572 before blaming the rule | issue #595 (duplicate half, @ai-dossier/sched >= 0.22.0), issue #613 (mis-attribution half, >= 0.22.2) — see docs/reports/batch-cycles-checkpoint.md §4.8 @ai-dossier/sched >= 0.48.0 (#810): evictions= counts only kind: evicted records (a member’s own hand-back is kind: handed-back and never counts); already-parked members show under preserved=, in-flight members under parked=, and a full dissolve with validated members becomes batch-blocked dissolve-refused:<why> instead. |
two eviction records naming the SAME member minutes apart while a LATER member of the same batch has no record at all, and executing_member jumped by two (serial batches only — under #809 parallel dispatch executing_member is the highest dispatched member index and jumps by design) | Distinct from the eviction-duplicate row above: that one is a duplicate ARRAY entry with correct attribution; this is MIS-attribution plus a skipped member. Two scheduler ticks each read executing_member before either one’s advance landed, so both resolved the SAME member — the second advanced the pointer a second time, past a member that never ran and never got a record. Do NOT reach for the dedupe: de-duplicating an append cannot make a record name the right member | Fixed (#613). The one-shot claim in advanceMemberOrValidate re-reads b.executing_member !== currentMember under the lock and no-ops on mismatch, and evictMemberDirectly journals unit-failed itself behind that same claim. Before blaming attribution logic, check the deployed @ai-dossier/sched carries #613 (>= 0.22.2), and look for member-advance-skipped in events.jsonl — a lost claim is journaled, not silent. Note a tick-driven test does NOT reproduce this (each tick re-derives the member from fresh state); only two resolutions against ONE captured snapshot do | issue #613; docs/reports/batch-cycles-checkpoint.md §4.8 |
A batch member requeued mode=full, batch=null, dispatch_profile=null after a hand-back (status=blocked reason=<x>), an agent-exited-unverified, or a dissolve — and restarted from the base branch on the config-default provider; or an already-VALIDATED member requeued when its batch dissolved | Pre-#810 the direct eviction rail always called requeueMember(... { mode: 'full' }), requeueMember never carried the batch’s dispatch_profile (a slot member’s own is null — the BATCH holds it), and checkDissolveTrigger counted a member’s own hand-back as an eviction (imboard b-20260924-02: 2 “evictions” > threshold 1 → dissolve → validated #4137 requeued) | Fixed (#810, @ai-dossier/sched 0.48.0): members PARK (evicted / handed-back) with their branch + profile; only kind: evicted counts toward the threshold; a threshold crossing with validated members is suppressed (dissolve-suppressed) and a full dissolve over them blocks (dissolve-refused:<why>). Remedy for a parked member: ai-dossier sched requeue --issue <n> (listed under == Parked members == in sched status). On an older engine, stop the requeued entry (sched stop --issue <n>) before its next tick and re-run it by hand from the member branch. #840 (@ai-dossier/sched 0.57.0): an aggregate-suite (post-landing) eviction parks too — a landed member’s remote branch now survives until the batch ends — a PR-conflict split never re-batches a validated member, and a requeued member’s first dispatch seeds a setup done milestone on its branch (resume-seeded) so the gate resumes ON it; resume-seed-failed in the journal means the run fell back to prompt text only | |
a batch TAIL respawned every ~2 min after posting batch-review status=blocked reason=members-mismatch (journal: unit-failed reason=tail-agent-exited-unverified with empty detail, repeating), or any tail/report agent that keeps exiting without a verdict | Pre-#832 the tail’s {members} was the raw batch.members, which still listed a member that had handed back and been sched stop-ed (no boundary commit on the integration branch → the tail correctly refused); the engine read the tail’s own blocked milestone as an unverified exit, and nothing capped the runBatchTick respawn wedge (imboard b-20260924-04: 4 strong-tier respawns, ~375k tokens) | Fixed (#832, @ai-dossier/sched 0.51.0): the tail gets the LANDED (validated) members, pre-checked both ways against batch.ranges (no-landed-members / members-mismatch:* block without spending an agent); a tail blocked milestone blocks the batch once as tail-blocked:<reason>; unverified exits are capped per phase (MAX_BATCH_AGENT_RESPAWNS = 2 → respawn-cap:tail, or a merged batch closes without its report); sched stop --issue accepts a parked member. Grep batch-blocked / unverified exit n/3 in events.jsonl; sched status prints the remedy | Issue #832 |
ai-dossier runstate post (or plan post) printed ⚠️ No milestone on issue #<n> carries run r-<n>-… and the ✅ link points at a DIFFERENT repo (e.g. imboard-monorepo/issues/<n> while you work on ai-dossier#<n>) | The CLI resolves the target repo from the CURRENT DIRECTORY’s git remote; an agent’s shell cwd resets between calls, so a post run outside the intended repo lands on the same-numbered issue of whatever repo the cwd is in (observed during #832’s run: a ship awaiting-merge milestone on imboard-monorepo#832) | cd <repo worktree> && in the SAME command as every runstate post/plan post, and read the repo in the ✅ URL before moving on; delete a misdirected post with gh api -X DELETE repos/<owner>/<wrong-repo>/issues/comments/<id> and repost from the right cwd | Issue #822 |
a batch member evicted agent-exited-unverified (or wrong-procedure) whose issue trail shows full-cycle milestones — review done next=ship, wip(review) [skip ci] commits, no batch=/review= keys | The member ran the WRONG PROCEDURE (full-cycle instead of member-cycle; imboard #4174, mechanical-tier model) | Fixed (#822, @ai-dossier/sched 0.52.0): detected by isWrongProcedureMilestone (hand-backs excluded); the member is re-prompted ONCE in place (member-reprompted in events.jsonl, reprompted_members on the batch), and only a wrong trail posted after the re-prompt evicts it wrong-procedure (parked, sched requeue continues it); a stray run that already reached ship is evicted at once as wrong-procedure-shipped with its PR named (stray_pr) — close or reconcile that PR by hand. @ai-dossier/sched >= 0.56.0 (#844): the “stop the agent first” wait no longer hangs on an agent that ignores SIGTERM — 120 s (KILL_ESCALATION_MS) after the first SIGTERM (member-stop-requested) it is SIGKILLed through its process group and kill-escalated is journaled once; on older versions a wrong-procedure member that never resolves with a live pid needs a manual kill -9 -<pid> | Issue #822, #844 |
engine.test.ts 60-90 failures only when two make test runs overlap; agent-flag.test.ts EACCES on config save | Tests shared host state: engine.test.ts’s beforeEach deleted every /tmp/sched-engine-* dir (including a concurrent run’s live dirs), and CLI/core/sched tests resolved os.homedir() to the real ~/.dossier. | Every vitest config has globalSetup: test-support/isolated-home.mjs (per-run mkdtemp HOME); never sweep tmpdirs by name prefix — clean only dirs you created. | #896 |
| a push-triggered workflow (publish-packages, a deploy) did NOT run after a PR merged, with no failed check anywhere — and git log origin/main -1 --format=%B shows [skip ci] inside the SQUASH commit body | ship-issue’s Step 2.5 gate clears skip markers off the PR head, which is what decides whether pull_request CI runs. It does nothing about the squash-merge commit: GitHub’s default squash body is the concatenated list of the branch’s commit subjects, so every wip(<phase>): … [skip ci] commit the WIP Sync Rule requires gets re-imported into the message that lands on main — and GitHub evaluates the marker against that message, suppressing the push event entirely. The PR itself is green and merged, so nothing looks wrong: the publish or deploy simply never happened. Observed on PR #602 (909f467): the three preceding pushes to main all ran Publish Packages to npm, this one ran nothing. Harmless there (docs-only, no publishable source changed) but on a PR touching a package it is a silently skipped npm release | Pass an explicit body at merge time so the WIP subjects are never re-imported — gh pr merge <n> --squash --body "$(gh pr view <n> --json body --jq .body)" (or --body-file). After ANY merge, verify the base-branch push actually fired: git log origin/main -1 --format=%B \| grep -Ei '\[(skip[ -]ci\|ci[ -]skip\|no[ -]ci)\]' must find nothing, and gh run list --branch main --limit 3 must show the expected push-triggered runs for the merge sha. To recover a skipped release, re-run the workflow by hand (gh workflow run publish-packages.yml) — do not push an empty commit to main. One more turn of the screw: a commit message that describes this trap and spells the marker literally triggers it too — the commit adding this very row did, and had to be reworded. Write it as “a CI skip marker” in commit messages; only the file content may carry the literal token. Structural fix (imboard-ai/git/ship-issue@1.15.0, #747): the per-issue merge step now calls the REST merge endpoint (gh api -X PUT .../pulls/<n>/merge) with an explicit, marker-free commit_message — the same Closes #<issue_number> body Step 6 already passes, never the PR body — so the default squash body’s concatenated wip(...) subjects (and anything else a PR body might contain) can never reach the base-branch commit — this mitigation is now structural, not agent discipline about passing --body | PR #602 (#598); structural fix #747 |
| claude -p dispatch log holds ONLY the two sched-dispatch preambles (no init event), the child stays alive burning CPU, and dispatchLogPath never grows | The #685 default --settings PreToolUse background-guard hook hangs under claude 2.1.266 — when the CLI invokes the guard’s node -e command the hook never returns, so the headless session emits zero stream-json events and the dispatch wedges until killed (verified with the identical argv: guard attached = silent hang, --settings {"disableAllHooks":true} = normal session) | Give the dispatch template its own --settings (an operator template carrying --settings suppresses withBackgroundGuard’s injection by design, dispatch.ts): dispatch.command: [claude, -p, ..., --settings, '{"disableAllHooks":true}'] — and verify the claude version on dispatch hosts before wiring #685-style guards | PR #705 (found during #677’s AC4/AC5 real-dispatch verification) |
| a gate/CI cost distribution says the gate is cheap, but a full run visibly takes an hour | Cost stats over gate runs were computed across ALL runs, including ones that failed at an early gate and exited before the expensive stage. Measuring imboard’s scripts/ci-parity.sh this way gave a median of 7.4 min over 124 runs — which reads as “ci-parity is cheap, batching is not worth building”. Splitting by outcome inverted it: green runs that actually reached the integration stage cost a median of 52.5 min (p90 89, max 120), and the short runs were early hygiene-gate failures. Two runs over the same 13-file diff produced 0 and 107 integration suites respectively | Split every gate-cost distribution by outcome AND by whether the run reached the stage being priced, before quoting a median. For ci-parity specifically: green ⇔ no [FAIL] gate line in the log, and “reached the expensive stage” ⇔ the log names >50 integration:<suite> targets. Durations come from worktrees/*/.ci-parity/run-<ISO>Z-<pid>.log — start time in the filename, end time from mtime | RFC-0001 §J.12 |
| local test run is green but CI fails on the same commit / a stack trace names a line number your patch changed | The fix was verified against an uncommitted working tree. The local run picked it up from disk; the pushed branch never had it. Cost a full CI cycle on imboard#4113 | Commit and push BEFORE verifying, or verify the committed tree explicitly (git show origin/<branch>:<path>). The fast diagnostic: if the stack trace cites a line number that your patch moved, the running code does not contain your patch — it is not an environment difference | RFC-0001 §J.16 |
| a Monitor stays silent through a real failure, or reports a still-running job as finished | Four distinct instances in one session, all one shape — a signal read as meaning something it does not. pgrep -f "<script>.sh" matches the monitor’s own command line, so its liveness check can never go false; a log quiet for 150s was a jest suite running --silent at 25%; a filter on [FAIL] gate missed the runner’s actual ✗ ci-parity: gate '<x>' failed; gh pr checks exits non-zero when anything is pending, which is a status not an error | Prove liveness by captured pid (/proc/<pid>), never by pattern match. Require an explicit terminal marker from the runner — silence is not completion. Enumerate every failure spelling the runner emits. Read a tool’s exit-code contract before treating non-zero as failure | RFC-0001 §J.16 |
| a blind Conformance re-run’s git diff <base_branch>...HEAD shows dozens of unrelated files across the whole repo, when the issue’s own Scope section names one or two | Run from the shared main worktree, <base_branch> resolves to the local main ref, which nothing fast-forwards — other fleet runs merge into origin/main continuously while this worktree’s own main branch sits behind. git diff main...HEAD then diffs against that stale ref, so the “diff” a Conformance re-run reads includes every commit that landed on origin/main since the worktree was last synced, not just the PR’s own changes. Confirmed on issue #654’s ship-issue Step 3a.5 re-verification: git log --oneline main..HEAD showed 13 commits (10 foreign, from unrelated merged PRs) vs git log --oneline origin/main..HEAD’s 3 (all this PR’s own) | Before diffing for any Conformance/Agent-7-style re-run (ship-issue Step 3a.5, review-issue Agent 7) done from a shared worktree, git fetch origin <base_branch> --quiet and diff against origin/<base_branch>...HEAD — never a bare local branch name, which cannot be trusted fresh in a worktree other runs don’t keep synced | issue #654 ship re-verification |
| Base branch was modified. Review and try the merge again. (GraphQL, from gh pr merge --squash --match-head-commit) | A different fleet tail run merged its own PR into the same base branch in the window between ship-issue Step 6’s mergeability check and the merge call — a race on the BASE ref, not this PR’s head. gh pr view <pr-number> --json mergeStateStatus flips to UNKNOWN immediately after while GitHub recomputes against the new base tip; it is not a --match-head-commit mismatch (this PR’s own head never moved) and not a real conflict | Transient, not a hard blocker: wait ~15–30s for mergeStateStatus to leave UNKNOWN, confirm it settles back to CLEAN/UNSTABLE (the same states Step 4’s gate already tolerates), then re-run the IDENTICAL gh pr merge --match-head-commit <PR_HEAD> … command unchanged — the flag still protects against this PR’s own head moving. Do not drop --match-head-commit to work around it, and do not spend a Step 5 CI-fix attempt on it — the PR’s own checks are unaffected. Structural fix (imboard-ai/git/ship-issue@1.15.0, #747): the per-issue merge moved to the REST endpoint and this case now surfaces as a 405 with Base branch was modified from gh api -X PUT .../pulls/<n>/merge; the fix is the same wait-and-retry, re-running the identical call (still keyed on sha=, the REST equivalent of --match-head-commit) unchanged | issue #655 ship re-verification (concurrent with #652/#653/#654); structural fix #747 |
| gh pr checks shows test (20)/test (22) blank or cancelled and mergeStateStatus=UNSTABLE, while a LATER CI run on the same head sha is green | A superseded CI run’s cancelled check-runs stay attached to the head sha alongside the newer, successful run’s check-runs — gh pr checks (and a naive “wait for this check-run to finish”) can land on the stale, cancelled one and stall waiting for a status that will never change. Confirmed on d1ed452: run 34328570742 (cancelled) and run 34328632358 (success) both attached to the same head sha; two tail agents each lost ~90 minutes waiting on the cancelled run before finding the green one | Evaluate the LATEST check-run per name, not the first one returned: gh api repos/<owner>/<repo>/commits/<sha>/check-runs, group by check name, and take the run with the latest started_at for each name. Never wait on a cancelled check-run to change state — it won’t. Record ci_note=superseded-cancelled-checkruns when this is what happened | issue #657 (PR #671) |
| npm ci fails Missing: @ai-dossier/sched@0.24.0 from lock file on every job of a PR that bumped a 0.x workspace version | A PR bumped packages/sched 0.24.0 → 0.25.0 but synced neither the root lockfile nor consumer ranges. The bare “run npm install to sync the lockfile” remedy is a trap here: on 0.x, caret pins the MINOR (^0.24.0 = >=0.24.0 <0.25.0), so after the bump cli’s range no longer admits the workspace package — npm “fixes” the sync by adding a registry-pinned published @ai-dossier/sched@0.24.0 under cli/node_modules/ (and npm ci then wants exactly that entry). The lockfile turns green while the monorepo silently decouples: cli’s imports resolve to the stale npm copy, so the sched fix under review is never exercised by cli code or tests | Bump every consumer’s range in lockstep (cli/package.json ^0.24.0 → ^0.25.0 — the #634 precedent), then npm install --package-lock-only to regenerate: the correct diff is exactly two lines (the consumer’s range in the workspace entry + the bumped package’s version). Prove the link with node -e "console.log(require(require.resolve('@ai-dossier/sched/package.json',{paths:['./cli']})).version)" → must print the NEW version and resolve into packages/sched, then npm ci, build, and test before pushing | issue #679 (PR #687 CI repair) |
| a batch member’s test.focused gate times out at the 900s cap and the batch blocks automation-broken — most reproducibly when the member touches two workspace roots | CapOutcome derives from the exit code alone, so “the capability ran past its duration cap” was indistinguishable from “the harness is broken” — a statement about DURATION blocked the whole batch. The timeout is real but earned by the selection: the member touched two workspace roots, so the gate resolved 3 package roots plus their closure instead of the “focused” set its name implies. #594’s earned-evidence branch does not cover it (the failure evidence is genuine — it is just about time, not reliability), and #585’s inconclusive-blocks path gives no skip route either. The resolved package list and duration were not recorded either, so operators could not see what “focused” actually selected | Fixed (@ai-dossier/sched >= 0.27.0, #681/PR #690). A gate TIMEOUT is declined, not blocked: journaled capability-unavailable (member skipped — the batch parent’s expensive aggregate stage covers it) instead of automation-broken; the resolved package selection and duration_ms are recorded per change shape for gate-cost reporting; a sched resume --batch <id> recheck dissolves an existing block whose recheck classifies as a timeout. Gate taxonomy now: inconclusive → block (#585); unevidenced task-failed → block (#594); timeout → decline (#690) | PR #690 (#681) |
| a publish-gating make test run times out one test at exactly 5000ms, and a rerun of the identical commit passes | Vitest’s DEFAULT testTimeout (5000ms) with no override in cli/vitest.config.ts. The sched auto-upgrade tests do real tmpdir fs work per run — ~1s locally (38-run distribution 926–1322ms, unimodal) — so the margin to 5s is thin on a shared, throttled CI runner that has already run lint + every other workspace’s suite. It reads as an intermittent product bug (“engine hang?”) but is fs-latency pressure; the tempting wrong fixes are re-running the flaky test in CI or debugging the mocked engine path, which cannot hang (no real spawns). The rerun-passes signature plus a tight local duration distribution is the evidence that it is the NUMBER, not the code | Calibrate, then raise the suite-level testTimeout with the measurement recorded in a config comment (#696: cli at 30s; packages/worktree-pool already used 60s). A test that needs more than the raised number consistently is a genuine hang — investigate it instead of raising again. --repeat does not exist in vitest v4; loop the single test with -t "<name>" in bash and strip ANSI before parsing the per-test ms | #696 |
| Version-bump check passed but the release never appears on npm | scripts/check-version-bumps.mjs compares HEAD against the merge base, not the base-branch tip. A branch cut before another PR bumped the same package sees the old version, so bumping to a number that has ALREADY been published and merged passes the version-bump job green — and publish-packages then found that version already on npm (before #826 it silently skipped it; now publish-guard.mjs fails the run as a version collision), so the change merges and is never released | Before trusting the check, compare against the tip too: git fetch origin main && git show origin/main:<pkg>/package.json \| jq -r .version and npm view <pkg> version. If either equals your bump, merge origin/main and bump again. Now structural (#826): the check also compares against the base-branch tip and reports STALE (... is not above <tip> on the base-branch tip) — but only when CI runs after main moved; a PR that was green before the other merged is caught at publish time instead (the publish-packages.yml … needs publishing row, publish-guard.mjs) | PR #577 (#566); tip check #826 |
| batch PR remains CONFLICTING or has auto-merge-blocked after validated members landed | The PR-watch rail previously only waited for a merge and never entered recovery, leaving an integration branch stale even though handlePrConflict already existed. Treating recovery as a fresh shipment also risks replacing the pull request and re-batching validated members. | Route either state through handlePrConflict; validate that the integration ref belongs to the batch, rebase against FETCH_HEAD, force-push with lease protection, and return to awaiting-merge while retaining the existing PR. | PR #872 (#867) |
| Formatter would have printed the following content on a generated .json | Biome formats committed JSON, and JSON.stringify(value, null, 2) expands short arrays that Biome collapses ("agentClis": [\n "claude"\n] vs ["claude"]). A generator that writes a committed JSON artifact therefore produces a file failing make check on every regeneration, silently fixed once by hand and then re-broken by the next run — including every scheduled regeneration PR | Exempt generated data from the formatter rather than formatting it by hand: add a biome.json override with "includes": ["<generated dir>/**"], "formatter": { "enabled": false }. Do NOT couple the generator to Biome | PR #577 (#566) |
| gh pr view <branch> reports a PR that is not open | gh pr view <branch> --json number resolves CLOSED and MERGED PRs for that branch, not just open ones. A recurring job that used it as an “does a PR already exist?” test kept editing the long-merged PR after the first merge, confirmed success, and left every later refresh pushed to a branch with no open PR for anyone to review | Ask for open PRs explicitly: gh pr list --head <branch> --state open --json number --jq '.[0].number // empty', create when empty, and re-read after create/edit to confirm | PR #577 (#566) |
| gh pr merge --squash returns 502 Bad Gateway, then every retry returns GraphQL: Merge already in progress (mergePullRequest) while gh pr view --json mergedAt keeps showing null | GitHub’s GraphQL mergePullRequest mutation can accept and start processing a merge, then time out the HTTP response with a 502 before the client sees success. The merge stays server-side “in progress” — every subsequent gh pr merge retry (also GraphQL) hits the same stuck lock and errors instead of completing or clearing it. Polling mergedAt never turns non-null because the GraphQL path is the one that is stuck, not the merge itself | Retry via the REST endpoint instead of gh pr merge (which always uses GraphQL): gh api -X PUT repos/{owner}/{repo}/pulls/<pr>/merge -f merge_method=squash -f commit_title="..." -f commit_message="..." -f sha=<PR_HEAD> — this independently completed the merge ("merged": true) when 4 GraphQL retries over ~2 minutes all returned the stuck-lock error. Always keep --match-head-commit/-f sha= on whichever path you use, so a push landing mid-retry still aborts instead of merging silently. Structural fix (imboard-ai/git/ship-issue@1.15.0, #747): Step 6 now calls the REST endpoint directly (not as a manual workaround) with a bounded poll-then-retry loop around any 5xx. One difference from the GraphQL shape this row diagnoses: there, mergedAt stayed null because the merge itself was stuck; on the REST path a 5xx can instead mask a merge that actually completed, so the loop polls mergedAt first and only retries the identical PUT with backoff if it is still null, at most 5 attempts over ~3 minutes, then reason=merge-unavailable — so the agent no longer needs to remember to switch off gh pr merge under pressure | PR #744 (#743); structural fix #747 |
| npm view @ai-dossier/<pkg> version (and npm i -g …@latest) still report the OLD version for 1–3 minutes after publish-packages.yml logged + @ai-dossier/<pkg>@<new> | The npm registry’s packument/dist-tag read path is eventually consistent and the local npm cache also serves the stale packument; the publish itself succeeded. Treating the stale read as “publish did not happen” invites a needless re-dispatch of publish-packages.yml | Confirm from the run log (gh run view <id> --log \| grep '+ @ai-dossier/'), then poll curl -s https://registry.npmjs.org/@ai-dossier%2f<pkg> \| jq -r '."dist-tags".latest' until it moves; install with the exact version and --prefer-online (npm i -g @ai-dossier/cli@<new> --prefer-online) | PR #806 (#804) |
| git stash pop applies a diff that touches files you never edited | .git is shared across EVERY worktree of a repo (refs/stash is one ref, repo-wide, not per-worktree). Two concurrent agents each running git stash push then git stash pop around the same moment race on that one ref: an agent’s pop (no index = top of the list) can grab a DIFFERENT agent’s just-pushed entry instead of its own, silently applying (and dropping) someone else’s uncommitted work into your worktree. It surfaced while diffing #789’s change against a clean baseline (git stash push -- packages/sched/ then git stash pop) and picked up an unrelated #791 run’s WIP on cli/status.ts/status.test.ts instead The same race had a second, silent path: .husky/pre-commit ran npx lint-staged, and lint-staged 15 backs up partially staged work with git stash by default — so every commit through the hook touched the shared stash ref even when agents obeyed the no-stash rule (#853). | Never use git stash in a repo with concurrent worktree agents — diff against a clean baseline with a saved patch file (git diff > /tmp/x.patch, or copy the file, edit, restore) instead, never the shared stash ref. If it already happened: check git stash list first — your own entry may still be sitting there untouched if only YOUR pop misfired (the other agent’s stash is what you have to hunt for). git fsck --unreachable --no-reflogs finds every dropped stash commit still in the object database (their content isn’t gone, only the ref pointing at it — yours too, but only if the other agent ALSO popped) — git log -1 --format=%s <sha> on each to identify whose is whose, git stash store -m "<original message>" <their-sha> to hand the other agent’s entry back onto the shared list, then git stash apply <your-sha> (never pop, to avoid re-touching the shared ref) to recover your own. The hook now runs npx lint-staged --no-stash (#853), so a commit through it no longer touches refs/stash. The cost: --no-stash implies --no-hide-partially-staged and skips the revert-on-failure backup, so unstaged hunks in a partially staged file are formatted and committed along with the staged ones. Agents stage whole files, so that is acceptable here; if you stage by hunk (git add -p), commit with the other hunks saved elsewhere first. | #789; hook path #853 |
| CI lint job (or the publish workflow’s Lint step) red with Some warnings were emitted while running checks on a Biome warning (e.g. lint/style/noNonNullAssertion, lint/suspicious/noExplicitAny, lint/suspicious/noTemplateCurlyInString), although a local biome check passed | Since #859 npm run lint (CI’s lint job, the publish Lint step, make lint) is biome ci --error-on-warnings ., and npm run check / make check / the pre-commit hook also fail on a remaining warning. A bare biome check or npx biome check --write without the flag still exits 0 when only warnings remain, so running Biome directly hides them. Info-level diagnostics fail none of these | Reproduce with make lint, never a bare biome check. Fix the warning; where the rule is wrong for that site (a literal ${{ … }} Actions expression or a shell ${VAR:-} inside a JS string), escape it in a template literal (`\${{ … }}`) or suppress it with // biome-ignore lint/<group>/<rule>: <why> | #859 |
Fleet / host ops traps
Traps from operating the autonomous fleet on its host (scheduler state, cron, process management) rather than from changing code in this repo. Same table shape; grep it the same way.
| Symptom (grep this) | What went wrong | Fix | PR |
|---|---|---|---|
cron job stopped firing after a script was rewritten / Permission denied from a cron line that used to work / a script rewritten with os.replace lost its +x | A Python helper rewrote an executable script with the write-temp-then-os.replace() pattern. os.replace swaps the inode, so the replacement carries the temp file’s mode, not the target’s. That mode is umask- and API-dependent, never the original: tempfile.mkstemp() gives 0600, a plain open() gives 0666 & ~umask (0664 on this umask-002 host) — either way the +x bit is gone. cron then fails to execute it, silently: cron’s own failure mail goes nowhere on an unconfigured host, the script’s log file simply stops growing, and ls shows a normal-looking file with the right name, size and mtime. It cost a 3.5 h fleet outage — the tick loop was dead the whole time and nothing said so | Copy the mode onto the TEMP file before the rename — shutil.copymode(path, tmp) then os.replace(tmp, path) — so the rename publishes a file that already carries the right bits. The stat-before / os.chmod-after variant works but is NOT atomic: between the replace and the chmod the file is live at the target path with the temp’s mode, which for a 0600 secret like ~/.dossier/reset-fleet/telegram.env means a real exposure window. Verify with ls -l on the target path, never on the temp. When a cron job goes quiet with no error anywhere, check the exec bit before you check the code | reset-fleet incident |
sched abandon returned success but the agent is still running / a slot’s work keeps appearing after you abandoned it | ai-dossier sched abandon --issue <n> (and --batch <id>) is a state-machine operation: it marks the unit failed and releases the slot so the queue can move on. It does not signal the agent process that unit spawned. The agent keeps running with the same worktree checked out — still committing, still pushing, still able to open a PR — while the scheduler believes that slot is free and dispatches something else into it. Two agents on one worktree is how you get a diff nobody planned | Treat abandon as “stop tracking this”, never as “stop this”. To actually stop the work: read the pid from sched status’s Slots table, then prove it is still that agent before killing — ps -p <pid> -o pid=,lstart=,args= must name the expected issue or worktree and have started before the unit’s dispatched_at (sched status pids go stale, see the row below, and a reused pid post-reboot belongs to something else entirely). Then kill <pid>, re-check ps -p <pid>, and only then abandon. If ps -p <pid> already prints nothing, the slot is merely stale — reconcile with ai-dossier sched start --once --project <slug>; do not go hunting for a replacement pid to kill | issue #472 |
your ssh session dies the instant you run pkill -f <pattern> on a remote host | pkill -f matches against the full command line of every process, including the shell running the pkill and the sshd/bash carrying the command you typed. Any pattern specific enough to name the agent you want (claude, full-cycle, the issue number, the worktree path) is also present verbatim in your own command line, so it matches your own session and kills it. The remote command is severed mid-flight, so you never see output and cannot tell whether the intended target died, whether anything died, or whether the kill ran at all | Never pkill -f over ssh. Resolve to pids first and dry-run: pgrep -A -af <pattern> — -A/--ignore-ancestors drops your own shell and its parents (it does NOT drop an unrelated peer session, so still read the list) — then kill <pid> by number. Two flag traps: -x combined with -f requires the pattern to equal the WHOLE command line, so adding it to a substring pattern silently matches nothing; and the age filter is -O <seconds> / --older <seconds> — --older-than is not a real option and exits with a usage error that reads like “no matches” | reset-fleet incident |
no claude / agent process visible in ps, so the unit is assumed dead and is redispatched or abandoned | “I see no process” is not evidence of death, and acting on it duplicates work. The agent may be a child under a different name (the CLI execs a node process whose argv does not contain the string you grepped), on a different tty or user, in a container, or momentarily between execs. The inverse also holds: sched status can show a slot running on a pid that is long dead — this host carried two such slots for three days (issue:3393 pid 3675888, issue:340 pid 3442401, both dead, both issues closed and shipped) because the project was paused and no reconcile tick ran. Neither direction of the process view is authoritative | Use the durable record, not ps: the issue’s runstate trail (ai-dossier runstate last --issue <n> --json), new commits on origin/<branch>, and the issue/PR state on GitHub. ps -p <pid> is meaningful only for the exact pid sched status recorded, and only about that pid. To clear slots stuck on dead pids run one reconcile pass (ai-dossier sched start --once --project <slug>) — it re-detects running slots by pid identity — rather than editing state.json. Run it while the project is still paused: paused gates dispatch only (state.paused ? [] : runnableUnits(...), packages/sched/src/engine.ts), never reconcile, so a paused pass frees the slots with no chance of spawning into them | issue #472 |
a sched status queue row reads done but the issue is still OPEN and carries no new runstate milestone | It never ran. done in the queue table is the scheduler’s own bookkeeping for “this unit left the queue”, and a unit can leave for reasons that produced no work — a verification that resolved against a stale trail from an earlier run, a tail/report unit that completed while the issue’s own work did not, a completion inferred before #576/#582 tightened it. Reading done as “shipped” is how a re-enqueued issue gets a second report posted for work nobody did | Cross-check every done against ground truth before believing it: the issue state on GitHub, a runstate milestone stamped after this dispatch (ai-dossier runstate verify --issue <n> --dispatched-at <iso> — that flag is what makes the check dispatch-relative), and a PR that actually merged. A done row with an OPEN issue and no post-dispatch milestone means the unit must be re-enqueued, not reported on | PR #586 (#582); PR #593 (#591) |
resume_from=ship-wait and the phase=ship status=awaiting-merge milestone carries no verdict_head=/verdict_refreshed=/verdict_check= keys | ship-issue’s Step 3a.5 verdict-freshness gate is defined to run before every merge authorization, but an awaiting-merge milestone posted by an older run (or one that skipped recording the keys) leaves a tail run resuming at ship-wait with nothing to “carry forward” at Step 6 — there is no prior result to compare the head against, so the “only re-run if the head moved” carve-out has no baseline. Superseded (ship-issue >= 1.13.3, #662/PR #668): the gate itself no longer compares heads by raw sha — it compares git patch-id --stable over the diff each head produces, so a rebase, squash, or a message-only rewrap (same content, new sha) reads as fresh instead of stale | Before merging, recompute the check yourself: pull VERDICT_HEAD from the last trusted phase=review status=done milestone’s head=, PR_HEAD from gh pr view <pr> --json headRefOid. Under verdict_check=patch-id (ship-issue >= 1.13.3), compare git patch-id --stable over git diff $(git merge-base origin/<base> VERDICT_HEAD)..VERDICT_HEAD against the same range ending at PR_HEAD — EQUAL patch-ids means fresh (no Conformance re-run) even when the shas differ; UNEQUAL means stale, re-run Agent 7 alone against PR_HEAD before authorizing the merge. Fall back to the old literal sha-equality test only when VERDICT_HEAD is unfetchable (verdict_check=sha-fallback) — record which comparison decided it | issue #652 (PR #658); patch-id gate: issue #662 (PR #668) |
nohup command’s log stops after its startup line while state.json/slot pids go stale — the “started” daemon is gone and nothing announced it | nohup <cmd> without a trailing & does not background anything — it blocks SIGHUP and NOTHING else. The command runs in the foreground of the agent’s tool call and dies with that call’s process group when the call ends. An agent “started” the sched engine this way (nohup ai-dossier sched start >> engine.log 2>&1): the engine ran exactly one tick, its log ended at the startup lines, slots stayed running in state.json for 14 minutes against a 60-second tick, and only comparing slot PIDs against /proc revealed the state was stale. The failure is easy to hit and hard to see: the name suggests detachment, the interactive form is nearly always typed with & so the &-less form reads correct to anyone skimming, and it is DELAYED (the process starts, logs its startup line, only disappears when the tool call ends) — so the agent’s own verification (“the log shows it started”) passes | nohup’s job is SIGHUP immunity; backgrounding is the &’s. Correct forms: nohup <cmd> >> log 2>&1 & — background AND SIGHUP-immune; setsid <cmd> >> log 2>&1 & — also leaves the caller’s process group; systemd-run --user --scope <cmd> — supervised. Then verify liveness AFTER the call returns, not at start: pgrep -f '<cmd>', or the log’s mtime advancing by more than one tick interval — a startup line is not proof of a running daemon. A long-lived engine from an agent tool call is the wrong deployment shape regardless of quoting: hold no daemon open inside a bash call — the fleet runs the engine as tick.sh + cron (sched start --once per tick, see docs/how-to/autonomous-pipeline.md) or a supervised unit (KillMode=process, packages/sched/README.md “Supervised deployment”) | issue #678 |
Evidence does not match dossier: ... does not match "<namespace>/<name>" from ai-dossier publish right after an evidence add that reported success | evidence add/evidence init default their sidecar’s dossier namespace exactly like publish does (--namespace, else first org, else username) — so an author who belongs to an org but is publishing this dossier under a personal namespace (or vice versa) gets a sidecar silently stamped with the default, not the namespace they actually publish to. evidence sync could not fix it — it only refreshed version/checksum, never dossier | Fixed (@ai-dossier/cli 0.48.0, #743). evidence sync --namespace <ns> now rewrites a mis-stamped dossier in place (same default chain as publish) — no need to delete and recreate the sidecar. A plain evidence add/evidence sync (no --namespace) instead PRESERVES the sidecar’s existing namespace, so fixing one is always a deliberate --namespace pass, never automatic. #921: a registry path with extra segments (imboard-ai/git/<name>, published via publish --namespace imboard-ai/git) is unknowable to evidence add, so publish now rewrites a sidecar whose ONLY difference is the namespace to the id it is publishing under (different name / stale version or checksum still fail), and evidence add says when a new sidecar fell back to the default namespace | #736, fix #743, #921 |
Dossier has no checksum; run 'ai-dossier checksum <file> --update' or sign it first from evidence add, immediately after following publish-dossier’s own Step 2 | publish-dossier’s Step 2 instructs deleting the checksum/signature objects on edit (sign regenerates them), but evidence add/evidence init required a checksum to already exist in frontmatter — the two steps’ own instructions conflicted when evidence recording was inserted between them | Fixed (@ai-dossier/cli 0.48.0, #743). evidence init/evidence add now compute the checksum from the dossier’s current body when frontmatter has none, instead of erroring — ai-dossier checksum --update is no longer needed as a workaround, and imboard-ai/meta/publish-dossier@1.1.1’s Step 2b drops that line | #736, fix #743 |
review-issue’s Step 3 review agents (unnamed, per its own instruction) each “run” for minutes then come back The subagent ended without delivering a report through SubagentHandback | This is a distinct failure from review-issue’s documented “named agent → mailbox” trap — no name: was passed here. Each agent completed its analysis but never called SubagentHandback before its turn ended, so the orchestrator’s synchronous Agent call returns with no findings and only an agentId. Resuming with SendMessage({to: agentId, ...}) sometimes itself fails with "the agent that spawned you is no longer running" if the orchestrator’s own turn already moved on — the agent then has no path back to its original caller and must be told to relay via SendMessage to a still-alive recipient (a teammate or the team lead) instead. Left unhandled, this reads as “review agents hung” and invites re-dispatching work that already finished | Never trust an Agent-tool result whose output is just an agentId with no findings as “still running” — it already ran to completion; resume it once via SendMessage({to: <agentId>, message: "deliver your report now"}). If that resume also reports the spawning agent is gone, have the resumed agent relay its findings via SendMessage to any other reachable teammate instead of retrying SubagentHandback — it will keep failing for the same reason. Batch all 7 resume requests in one turn (independent, no need to serialize) | issue #751 |
A read-only gh pr view <pr> --json mergedAt,state,mergeCommit (or any command combining it with an unrelated write like gh issue edit --remove-label) is denied with Permission for this action was denied by the Claude Code auto mode classifier. Reason: [Merge Without Review], even though the actual merge already succeeded via a separate gh api -X PUT .../pulls/<pr>/merge call moments earlier | The permission classifier pattern-matches the bash command TEXT, not what it actually does — a compound command whose string contains mergeCommit (or is chained with another command via &&) can trip a “merge” classification even when every sub-command is read-only. It is not blocking the merge itself (that already happened); it is blocking the verification read that comes after it, and can cost time second-guessing whether the merge silently failed | Split compound commands apart: run the label edit and the --json mergedAt,state / --json mergeCommit reads as separate, single-purpose Bash calls instead of one chained string. If a merge-confirmation read gets denied, first check whether the merge already went through (the prior gh api -X PUT .../merge call’s own response usually already has "merged": true and the sha) before assuming the merge failed | issue #750 |
publish-packages.yml’s “Check if @ai-dossier/<pkg> already published, skipping for EVERY package on a merge whose own PR genuinely bumped that package’s version, and the merged commit never appears on npm even though check-version-bumps.mjs passed green on the PR | This PR’s own merge-base comparison was correct in isolation — a DIFFERENT, unrelated PR independently bumped the same package to the same target number from the same starting version (both picked “next minor” from an identical base) and merged first. The publish workflow’s concurrency group (cancel-in-progress: false) serializes the two push runs, so the earlier-merged PR’s run publishes that version number to npm before this PR’s own run reaches its publish step — which then finds the number already taken and silently skips, exactly as designed for a genuine re-run, but here the skip is masking a real, unpublished release. Distinct from the existing “compares merge-base not tip” row (PR #577): each PR’s own version-bump check was individually correct against ITS OWN merge-base — the collision is only visible at publish time, after both merges have landed | Fixed structurally (#826). publish-packages.yml now decides per package with scripts/publish-guard.mjs: it reads the published version’s source commit (gitHead, from the registry HTTP API GET https://registry.npmjs.org/<name>/<version>) and diffs the package’s release-relevant paths (same definition as check-version-bumps.mjs) against HEAD. Same commit or no release-relevant diff -> skip (re-runs stay idempotent); a diff -> ::error title=<pkg>@<ver> version collision::, unaffected packages still publish (dependents of the colliding one are held), and the job fails at Fail on version collisions. If the guard cannot decide (non-404 registry answer after retries, missing/unfetchable gitHead) the package is recorded as unavailable — never a silent skip: unrelated packages still publish, its dependents are held, and the final step fails the job naming it. An unbumped src/ change merged under no-release-needed produces the same collision, by design. So the symptom is now a RED publish run, not a green skip. The fix is unchanged: a follow-up PR bumping the named package past that version (cd <pkg> && npm version patch --no-git-tag-version) — confirmed working: #823. Replay: the guard run against 97ddf58 (#820) with the real registry reports the collision against 85429e5 (#821) | issue #817 (PR #820); fix #826 |
publish-guard (packages/core) could not run / @ai-dossier/<pkg>@<ver> is on npm but has no usable gitHead () on the FIRST package of a publish run, although npm view <pkg>@<ver> gitHead on your machine prints a sha | publish-packages.yml installs npm@latest (for trusted publishing), so CI’s npm major moves under you. npm 12 prints npm view <pkg>@<exact-ver> <fields...> --json as a one-element ARRAY ([{"version":…,"gitHead":…}]) where npm ≤ 11 printed a plain object; a parser that reads .gitHead off the result gets undefined. The guard failed closed (exit 2), as designed, so nothing was skipped — but every publish run went red until fixed. A local npm 10/11 cannot reproduce it | Fixed structurally: publish-guard.mjs now reads gitHead from the registry HTTP API (GET https://registry.npmjs.org/<name>/<version>, a stable public manifest format) instead of npm view output; the npm CLI parser (accepting both shapes) is kept only as an inert fallback. And a package the guard cannot decide no longer halts the job: under --defer-collision it is recorded as unavailable <pkg>, unrelated packages still publish, its dependents are held, and --report-collisions fails the job naming it. When a CI step parses npm/CLI output, reproduce with the SAME major CI installs — npx -y npm@latest view … --json — not the local default | #826 follow-up (first publish after PR #839) |
A git stash pop you ran pops SOMEONE ELSE’S stash entry — your own working-tree changes vanish, replaced by an unrelated agent’s WIP diff on totally different files, even though you only ever ran git stash/git stash pop inside your own worktree | git stash is stored per-REPOSITORY (refs/stash, in the shared .git directory), not per-worktree — every worktree under the same repo shares one stash stack. Two concurrent agents in sibling worktrees (each running the ordinary git stash && git checkout <base> && test && git checkout - && git stash pop pre-existing-failure-check recipe, or a husky/lint-staged pre-commit hook’s own internal git stash backup) can interleave: agent B’s stash push lands on top of agent A’s between A’s push and pop, so A’s stash pop (which always pops index stash@{0}, the MOST RECENT entry) pops B’s stash into A’s working tree instead, and drops B’s entry from the shared list | Never use bare git stash/git stash pop in a repo where other worktrees may be active concurrently — use a patch file (git diff > /tmp/x.patch) or a scratch worktree instead. If it already happened: don’t panic-pop again. Diff the wrongly-applied changes against files you didn’t touch to identify whose they are, git checkout -- <those files> to revert them out of your tree, then git fsck --unreachable --no-reflog | grep commit and git show --stat <sha> on each candidate to find YOUR OWN stash commit (msg WIP on <your-branch>: ...) among the dangling commits — a git stash pop that “loses” an entry still leaves its commit object in the ODB, just unreferenced — then git checkout <that-sha> -- <your files> to recover it. husky/lint-staged’s own internal stash-backup step is subject to the exact same race and can silently succeed or interleave with a sibling agent’s — check for concurrent git stash/vitest/agent processes (ps aux | grep -iE "git stash|vitest") before any commit in a shared multi-worktree repo | issue #791 |
git patch-id --verbatim (ship-issue’s Step 3a.5 verdict-freshness gate) fails with a bare usage: git patch-id [--stable | --unstable] and no other output, breaking the freshness comparison silently (both VERDICT_PID/PR_PID end up empty, so the gate reads FRESH=false even when the diff is genuinely unchanged) | --verbatim was added to git patch-id in a newer git than what may be installed on the fleet host (confirmed reproducing on git 2.34.1, which only accepts --stable/--unstable) — the command exits non-zero with a bare usage line, not a recognizable “unknown option” message, so it’s easy to misread as a real diff-comparison failure rather than a flag-compatibility gap | Check git --version before trusting --verbatim; on git < 2.35ish, substitute --unstable (the direct predecessor flag — same “do not strip whitespace” semantics --verbatim documents, just the older name for it), never --stable (that strips whitespace, which is the exact failure mode --verbatim exists to avoid per the gate’s own comment). A usage: line with no other output from ANY git patch-id invocation in this gate means a flag-support gap, not a real staleness result — don’t let it silently degrade into treating every rebase as stale (extra unnecessary Agent 7 re-runs) or, worse, into skipping the gate | issue #791 |
A reconcileKeptWorktrees-style test that patches a done batch’s member_runs[] and expects your new evidence-based clearing to fire — instead torn_down becomes true (or, since #855, teardown_failed_at gets set) with ZERO of your expected journal events, plus a member-worktree-torn-down or teardown-failed/failed-invalid-worktree event you never wrote code to produce | runBatchTick’s terminal-batch safety net (TERMINAL_BATCH_STATUSES.has(batch.status), added by #809 for sched stop/abandon leaking member slots/trees) runs EARLIER in the same tick than reconcileKeptWorktrees, and it fires for EVERY terminal batch, not just the stop/abandon case. For a non-kept batch it calls teardownParallelRuns, a REAL teardown attempt (git worktree remove / pool return via the injected exec); since #855 it marks torn_down: true only when that attempt verifiably landed, and records teardown_failed_at (one teardown-failed line, never retried) when it did not — e.g. isSafeWorktree rejects a test’s non-realistic path. A kept batch (worktree_kept, set by the members-closed reconcile) is left alone. So by the time any code reachable only through runBatchTick inspects a NON-KEPT batch’s member_runs[], a live entry is usually already gone — any test that goes through the harness’s normal tick() observes the safety net’s effect, never your own function’s, and a naive assertion on “no destructive command” is also void: the safety net’s calls are real and unrelated to what you’re testing | Call the function under test directly and in isolation — export it test-only from batch-dispatch.ts (not the package’s public index.ts, same convention as reconcileStaleBlockedBatches/evictMemberAndContinue) and invoke it without going through runBatchTick. Reserve a tick()-based test for deliberately documenting the end-to-end interaction (assert what code path actually wins), not for testing your function’s own contract. Fixed in #855: member_runs[] now gets the same keep protection worktree/member_worktree get (persisted as BatchEntry.worktree_kept), and a failed teardown leaves torn_down: false — a test seeding a kept batch must set worktree_kept: true, and a test expecting a teardown must give it a realistic path (under <repoDir>/worktrees) and a cooperative exec/fsExists, or it gets teardown_failed_at instead | #834 (extending #791’s kept-worktree warning to member_runs[], #809); #855 (keep decision + verified torn_down) |
A CLI-level test does vi.mock('node:child_process') and asserts on a @ai-dossier/sched function’s gh/git calls (e.g. an anchor sweep or abandon-warning path), but the mock never intercepts anything — the real execFileSync runs (or, worse, silently no-ops in a sandboxed test env and the assertion just sees an empty call list) | @ai-dossier/sched is a workspace package the CLI imports across a package boundary; when a CLI test file pulls it in, vitest’s dependency-externalization behavior loads the package’s compiled dist/ via native require rather than through the test’s own module graph, so vi.mock on Node builtins set up in the CLI test never reaches exec calls made inside sched’s code. This is architectural, not a one-off mistake, and reads as “the mock silently isn’t working” rather than “this boundary can’t be mocked here”. The obvious-looking fix (test.server.deps.inline: ['@ai-dossier/sched']) does NOT work either, and fails the same silent way: it only changes whether Vite transforms the package’s already-resolved entry file, not what that file’s own require() calls resolve through — sched’s main is compiled CommonJS (dist/index.js), and per vitest’s own common-errors guide “Vite plugins, aliases, transforms, and module mocks do not apply to required files”, so an inlined-but-CJS sched still runs its internal execFileSync unmocked. Verified empirically on #829: inlining @ai-dossier/sched left vi.mocked(execFileSync).mock.calls empty after a call made from inside the package | Fixed (#829). cli/vitest.config.ts aliases the @ai-dossier/sched specifier straight to the package’s TypeScript source (resolve.alias: { '@ai-dossier/sched': '<repo>/packages/sched/src/index.ts' }), not its compiled dist/. Source uses real ESM import, which Vite transforms and resolves through its own module graph — including every module sched itself imports — so a mocked node:child_process reaches every exec call in the package, not just the CLI’s own. First consumer: the sched abandon --batch anchor-still-open end-to-end tests in cli/src/__tests__/commands/sched.test.ts (search #829). Caution for the next consumer of this alias: it also swaps in whatever the source currently does, which may differ from the last-built dist/ if someone forgot to rebuild — #829 hit exactly this: a sched start --once --auto-upgrade test that had silently relied on sched’s ground-truth gh/ai-dossier calls being unmocked (and thus failing outright in the sandbox, which read as “ground truth unreachable” and left a fabricated slot untouched) started flaking once those calls became real, working mocks that answered with an unrelated file-scoped default response instead of a deliberate one — fix by giving ground-truth-touching tests an explicit, scenario-appropriate execHandles stub rather than relying on the coarse “returns the same JSON to every call” default. This is a local-dev footgun only, not a CI/publish gap — verified: both .github/workflows/ci.yml’s test job (Build step, run: make build-all, immediately before Run tests with coverage) and .github/workflows/publish-packages.yml (make build-all at both its build and publish jobs, before npm pack/publish) always rebuild packages/sched/dist fresh from the exact commit’s src/ before it is either tested or shipped, so CI and the published @ai-dossier/cli tarball can never diverge from what this alias tested. The risk is scoped to a human/agent running CLI tests locally right after editing packages/sched/src without an intervening npm run build in that package — the CLI tests will reflect the edited source (correctly, via this alias) while a different local consumer of the compiled dist/ (a globally-linked ai-dossier binary, another package’s own build) would not yet | issue #790, fixed by #829 |
A merged version bump never reaches npm and gh run list -w publish-packages.yml shows NO run (not even skipped) for the merge SHA | GitHub intermittently drops the push event for a merge (no push-triggered workflow of any kind runs for that SHA, so it is not a paths filter or the publish concurrency group). Bursts of merges are the usual case; the last one is never rescued by a later push | Wait for publish-reconcile.yml (hourly; #798) to dispatch publish-packages.yml, or dispatch it yourself: gh workflow run publish-packages.yml --ref main. Check: node scripts/publish-reconcile.mjs --dry-run (needs GH_TOKEN) | #798 |
build-sea.mjs prints many warning: Can't find string offset for section name '.note' lines, or a macOS SEA binary is killed / postject output is corrupt | The warning: Can't find string offset lines come from postject’s LIEF on the stock Node ELF and are harmless (the binary runs). On macOS the copied node is already signed, so injecting without codesign --remove-signature first yields a corrupt binary, and an unsigned Mach-O is killed on Apple Silicon, hence the ad-hoc codesign --sign - after injection; vercel/pkg is archived, do not reach for it | Keep the remove-signature / inject / ad-hoc-sign order in scripts/build-sea.mjs; judge success by running scripts/smoke-binary.mjs <binary>, not by postject’s stderr. macos-13 runners are retired: the x64 macOS build uses macos-15-intel | #28 |
A batch blocks suite-unreadable with detail manifest declares gate.batch active but the capability layer reports it unavailable, on a repo whose test.full is timeout_prone: true | The manifest declares an active gate.batch, but the ai-dossier on PATH cannot see it (older CLI, stale shadow copy, not invocable). createBatchSuiteRunner used to fall through to cap run test.full — the timeout-prone gate the #777 enqueue refusal exists to keep batches away from — and the batch died on a gate that usually never produces a verdict | Fixed (#793). A declared-active gate.batch that comes back unavailable is now a terminal readable: false result (batch blocks, members preserved, test.full never run). Fix the install (ai-dossier --version, which -a ai-dossier, npm-prefix / stale-copy traps) and sched resume; do not “fix” it by un-declaring gate.batch | #793 |
Parallel batch evicts a member member-worktree-prep-failed:member-branch-checkout-failed right after an engine restart; the pool holds two claimed trees for one member | prepareMemberWorktree pool-claims a tree and re-points it at the member branch BEFORE the run is recorded. An engine exit in that window orphaned the claim at a pool path that is not re-derivable (the pool renames claims to the branch slug); the next tick claimed a second tree and checkout -B failed because the branch was checked out in the orphan | Fixed (#815). Prep first looks for a worktree already on the member branch (git worktree list --porcelain, orphanedMemberClaim) and takes it over (journaled member-claim-reused); no run was recorded, so no agent ever worked in it. Crash-injection test: batch-integration.test.ts “#815 item 1” | #815 |
Parallel batch blocks member-push-failed / landing-push-failed on a network blip | landParallelRun blocked the batch on the first failed push although re-landing is idempotent | Fixed (#815). A failed push is retried on the next MAX_LAND_PUSH_RETRIES (3) ticks (journaled landing-retry) before the batch blocks; the counter is in-memory, so an engine restart grants a fresh (still bounded) budget. Non-push landing failures still block at once | #815 |
VS Code extension bundle is megabytes larger than expected, or activate throws KMSClient/AWS errors | @ai-dossier/core imports @aws-sdk/client-kms at module top level, so bundling core into packages/vscode drags in the whole AWS SDK. | scripts/build.mjs aliases it to src/kms-stub.ts (throws on use); KMS signatures are reported as “use the CLI” by Verify. Do not remove the alias; keep the stub’s export names in sync with what packages/core/src/signature.ts and signers/kms.ts import. | #9 |
classify prescreen / batch compose / plan get warns Ignored N newer plan:v1 artifact(s) from author(s) who are not repo owner / org member / collaborator, plan validate says No trusted plan:v1 artifact, sched enqueue rejects no-plan-artifact/no-classify-record although the artifact is on the issue, or teardown ignores a setup milestone | plan:v1 and runstate:v1 artifacts are issue COMMENTS anyone can post. Before #808 the readers took the LATEST artifact by any author, so a stranger’s benign plan could dodge the file-count exclusion / path-floor review=full reasons, and slot-cycle could implement a different plan than the one screened. Every reader that ACTS on a comment now goes through isTrustedAuthorAssociation (@ai-dossier/core): GitHub author_association OWNER/MEMBER/COLLABORATOR only (a coarse project-membership signal, not proof of push access). There is NO BOT association: a plan/milestone posted with a GitHub App or Actions token reports NONE/CONTRIBUTOR and is NOT trusted. A missing or non-string association fails CLOSED everywhere (old gh) | Post from the operator’s own user (ai-dossier plan post / runstate post as OWNER/MEMBER/COLLABORATOR), and use a gh whose --json comments reports authorAssociation. Do not add BOT or widen the set to silence a warning. Readers: findLatestTrustedPlan/trustedCommentBodies (cli/src/plan-artifact.ts), parseSetupInfo (packages/sched/src/groundtruth.ts). #932: the scheduler’s milestone reads (latestMilestone/milestonesSince) pass runstate last/list --trusted, and runstate last/list/verify are trusted-only BY DEFAULT (--all = unfiltered operator reporting read; stats stays unfiltered); a stranger’s newer milestone is ignored with a stderr Ignored N newer runstate milestone(s) note. sched needs a CLI with --trusted (older CLI => unreachable, fail closed) | #808, #932 |
autoMergeRequest=null on a parked PR / awaiting-merge that never ends / no-merge-mechanism | The dispatch prompts hard-coded ship_mode=detached (park on the auto-merge label and STOP), but sched’s PR watch only WAITS — it never merges. In a repo with no watcher workflow where nothing ever REQUESTED native auto-merge (ai-dossier itself: allow_auto_merge=true, yet autoMergeRequest=null), the label is inert and the PR parks forever (#878; #869/#871 earlier). The dispatched run also never loaded ship-issue’s #860 merge-mechanism guard, because the prompt told it to park first | Fixed (#887/#874, @ai-dossier/sched >= 0.64.0). createExecGroundTruth().mergeMechanism() reads gh api repos/<r> (allow_auto_merge, allowed methods) and scans the REMOTE default branch’s workflows for a label watcher (10-min cache; Mergify/Kodiak, remote reusable workflows or an unreadable file => watcher unknown); the default full-cycle and batch-tail prompts render {ship_clause} from it — detached only when a mechanism is confirmed AND gh pr view --json autoMergeRequest is non-null after requesting, else attached (merge directly after review + green checks); a batch tail needs a confirmed WATCHER. A repo-level allow_auto_merge=true proves nothing about a PR: the backstop keys on the PR — reconcileParked fails a parked OPEN PR whose autoMergeRequested=false where no watcher exists (watcherWorkflow===false; unknown never trips it) as no-merge-mechanism once the condition persisted 10 min (no_merge_mechanism_since). An awaiting-merge milestone with ship_mode=attached is not a park (exit => redispatch). Not covered: the batch PR watch has no backstop, only the prompt. A custom dispatch.prompt without {ship_clause} keeps its own wording. Also: runstate refuses review done agents_done=prescreen (a non-reviewer), @ai-dossier/cli >= 0.79.0 | #887, #874 |
batch compose / classify prescreen hangs for tens of seconds on one issue body, CPU pegged in a regex | A new regex over an UNTRUSTED issue body with an unbounded nested quantifier and no anchor/lookbehind ((?:[\w.-]+\/)+… in a code-path test) is quadratic on a.a.a.… — 200 KB took 93 s. Existing scanners survive by lookbehinds or {1,N} bounds. | Bound every quantifier ({1,60}), cap the scanned body (GitHub’s limit is 65,536 chars), and time an adversarial string ('a.'.repeat(100000)) in a test. | PR #926 (#802) |
A respawned batch tail / member / takeover starts from the last PUSHED head and the gated-but-uncommitted files of the dead agent are gone (#920: 16 files, imboard b-20260929-01) | The engine respawned onto (or reset) a worktree without looking at what the dead agent left: only pushed commits survive a restart, and a gate.batch pass can describe code that was never committed | Fixed by #945: preserveWork (packages/sched/src/preserve.ts) commits tracked changes + secret-filtered, size-capped untracked files to refs/sched-rescue/<unit>-<ts> (a non-branch ref, 14-day TTL, pruned at sched start and daily) via a throwaway index, journals work-preserved, and the respawn prompt says to resume from it. Only tracked changes + unpushed commits are pushed; a rescue that includes untracked files stays local (local only), and pushed refs are world-readable on public repos and deleting a ref does NOT unpublish them (the TTL is housekeeping, not privacy); a tail respawn over uncommitted files or a newer dirty-tree gate row blocks tail-dirty-worktree (commit/discard, then sched resume --batch). Find lost work with git for-each-ref refs/sched-rescue and grep work-preserved events.jsonl. Needs @ai-dossier/sched >= 0.68.0 | #920, #940, #945 |
sched start log ends in ✓ … nothing to do and the engine is gone, no error, sched status shows stale-engine-lease | nothing to do is an idle tick’s line, not an exit path. Until #945 only SIGINT had a handler: SIGTERM/SIGHUP (another session’s restart, a closing terminal) killed the engine silently and left the lease stale; SIGKILL/OOM still leave no line | Look for the last engine-exit in events.jsonl (reason=signal:*, uncaught-exception + stack, …); none + stale lease = SIGKILL/OOM (journalctl -k). engine-restarted-after-crash / stale-engine-lease-alert events and sched status --alert make it visible; engine-hung flags a live pid with a stale heartbeat; SIGHUP is now ignored (journaled); run the engine under the supervised unit (#679) | #920, #945 |
A subcommand reading command.optsWithGlobals() (e.g. usage sync --since) silently gets the PARENT command’s default (usage --since = 30d), so its own “no flag given” cursor logic never runs and every run re-collects 30 days | Commander merges the parent’s options — defaults included — into optsWithGlobals(); a parent-level .option('--since', …, '30d') is indistinguishable from the user typing --since 30d. It only showed as collected 59220 rows since <now-30d> on a second, supposedly incremental, sync | Give shared parent options NO commander default and apply the default in the consumer (opts.since ?? '30d'). Verify an “incremental” path by running it twice and reading the since it prints, not by trusting the code path | #782 |
Rendered from docs/agent-traps.md in the repository. Edit it there.