AI Dossier

Capability Manifest & ai-dossier cap

Capabilities are a repo’s deterministic execution units — the recurring operations dossiers and workflows need (run focused tests, lint, build, install dependencies, prepare a worktree) expressed as named commands instead of re-reasoned by an agent on every use. Per the Progressive Determinism brief (RFC-0001), repos should accumulate these deterministic implementations and use them as the fast path, with reasoning as the fallback. The scheduler’s batch member gate consumes typecheck.run / test.focused from this manifest today (the per-member incremental gate, #583/#625).

  • Where it lives: .dossier/automation/manifest.yaml in the repo (resolved from the directory you run ai-dossier in).
  • Portability: a repo without .dossier/automation/ is a normal state — cap list prints an empty list and exits 0. Nothing breaks.
  • Preferred style: entries should mostly reference existing repo tooling — package scripts, Makefile targets — rather than duplicate logic:
# .dossier/automation/manifest.yaml
version: 1

capabilities:
  test.focused:
    command: npm test -- --silent
    lifecycle: active
    description: Focused vitest suite (fast path for agents)
    assumptions:
      - file-exists: package.json
      - tool-version: node>=20

  lint.run:
    command: npm run lint
    lifecycle: active
    description: Biome check

  test.full:
    command: make test
    lifecycle: shadow        # listed, but not executable yet
    description: Full suite incl. scripts — promote to active when trusted

Manifest schema

FieldTypeRequiredMeaning
version1noManifest format version (currently only 1; absent = 1)
capabilitiesmappingyesCapability id → entry
gatesnone-declared-on-purposenoDurable opt-out (#895) of sched enqueue’s undeclared-gate warning, for a repo that deliberately declares no typecheck.run / test.focused. The only accepted value; anything else makes the manifest invalid
entry .commandstringyesCommand line executed via the shell, in the directory ai-dossier runs in
entry .lifecycleactive | shadownoactive (default) = executable; shadow = declared but not yet trusted to run
entry .assumptionslist of probesnoPreconditions checked before the command runs
entry .descriptionstringnoWhat the capability does (shown by cap list)
entry .timeout_msnumbernoPer-entry command timeout in ms (default 5 min; a timeout is automation-broken)
entry .min_duration_msnumbernoSanity floor (#583): a non-zero exit that finishes faster than this is reclassified automation-broken instead of task-failed — “this probably didn’t really run”, not a genuine failure. Default 0 (no floor)
entry .timeout_pronebooleannoThe repo’s own admission (#777) that this capability routinely runs past any reasonable cap run budget. On test.full with no active gate.batch, sched enqueue refuses to form a batch (see the batch gate). Default false

Capability ids are dotted lowercase words (test.focused, worktree.prepare).

Assumption probes

Each assumption is a single-key YAML object. Probes run before exec; if any fails, the outcome is automation-broken and the command never runs.

ProbeExampleCheck
file-exists- file-exists: package.jsonPath (file or dir, relative to the run directory) exists
tool-version- tool-version: node>=20<tool> --version output satisfies <op><version> (ops: >= > <= < = ==; == is an alias of =)

cap init [--print] (#645)

Scaffolds .dossier/automation/manifest.yaml from detectProjectEnv (package manager, install/build commands) and the package.json scripts. Idempotent: an existing manifest is never modified. Undetectable gates become commented TODO stubs. sched enqueue warns (never blocks — #625) when a created batch’s manifest lacks typecheck.run / test.focused; opt out per invocation with --skip-gate-check / DOSSIER_SKIP_GATE_CHECK=1, or durably with gates: none-declared-on-purpose in the manifest (#895). With --repo <owner/name> the target repo’s manifest is read from --repo-dir <path>, or from the cwd when its origin remote is that repo; with neither, enqueue says the check was skipped rather than guessing a manifest.

cap list [--json]

Shows capabilities, lifecycle, command, and description. Absent .dossier/automation/ → empty list, success exit. A present-but-malformed manifest is a hard error (exit 1) with a message naming the problem.

cap run <id> [-- args]

Executes one capability. Extra args after -- are shell-quoted and appended to the entry’s command (cap run test.focused -- --grep auth → npm test -- --silent --grep auth). Args are data, not shell syntax — an arg containing ;, $(), or spaces reaches the command as a single literal word; put shell syntax in the manifest command itself. Only lifecycle: active entries execute; a shadow entry refuses with capability-unavailable.

The result is always one of exactly four outcomes, distinguishable by exit code and by a JSON envelope printed as the last stdout line (child output is passed through first; machine consumers should prefer the envelope file — see below):

OutcomeExit codeMeaning
ok0Command ran and exited 0
task-failed1Command ran and legitimately failed (e.g. red tests) — the operation’s failure, not the automation’s
automation-broken2Assumption probe failed · command missing/not executable (shell 126/127) · abnormal termination (signal / exit > 128) · timeout · manifest invalid · a non-zero exit faster than the entry’s min_duration_ms (#583 — “this probably didn’t really run”)
capability-unavailable3Id not in the manifest, no manifest at all, or lifecycle: shadow

The distinction matters to callers: task-failed means “trust the result — the task itself failed”; automation-broken means “do not trust the machinery — fall back to reasoning”; capability-unavailable means “no fast path here — reason from scratch”.

One caller qualifies the first of those, and it is worth knowing about before you write a capability script (#594): the scheduler’s per-member incremental gate only treats a task-failed as a red suite when output_tail carries evidence the run EARNED it — failing-test output for a test.* capability, compiler errors for the others. A capability that exits non-zero having produced no such output is routed to the block-the-batch path instead of evicting a member, because a wrapper that fabricates its exit code is otherwise indistinguishable from a genuinely red suite. Practical consequence: make sure a failing capability lets its runner’s own output through, rather than swallowing it and printing only its own framing.

Exit 1 is also the CLI’s generic usage-error exit (e.g. a typo’d command). Machine consumers should read the envelope — from --envelope-file / $DOSSIER_CAP_ENVELOPE_FILE, else the last stdout line carrying "cap_envelope": 1 — present for every cap run outcome, rather than the exit code alone, and check stderr for usage errors.

Machine consumers should read the envelope file, not stdout (#811). Pass --envelope-file <path> (or set DOSSIER_CAP_ENVELOPE_FILE=<path> in cap run’s environment) and cap run writes the same envelope there, atomically, before it exits. Stdout is a shared, lossy channel: a descendant process that writes after cap run prints the envelope pushes it off the last line, and (before #811’s stdout drain) process.exit() could drop a large re-emitted output’s tail — envelope included — which is how a green 33-minute gate.batch was once recorded as “no envelope”. The variable is stripped from the environment of everything cap run spawns for the capability (its assumption probes and its command), so the command is never handed the envelope path. Every envelope carries "cap_envelope": 1; a consumer without the file should scan stdout bottom-up for the last line carrying that marker rather than trusting the last line blindly — and, from either source, trust an envelope only when its outcome agrees with cap run’s exit code (the one signal the command cannot forge). The scheduler’s suite and per-member gate runners do all of this through cli/src/cap-envelope.ts.

On any non-ok outcome, the envelope also carries output_tail (#583 AC1/AC3) — the last --tail-bytes (default 8192) bytes of the command’s combined stdout+stderr, UTF-8-safe (never splits a multi-byte character). Omitted entirely on ok, so a passing run’s envelope stays small. The batch engine’s incremental gate uses this for attribution — a per-gate log file under the project’s runs/ directory and the last ~500 bytes in the journal unit-failed/gate-inconclusive event detail — rather than requiring a human to grep the raw agent transcript to find out why a gate blocked or evicted a member.

Capturing this output changes how cap run behaves for a human running it directly: output is now buffered and re-emitted after the command finishes, rather than streamed live via stdio: 'inherit' as before #583 — a long-running command shows nothing until it completes, instead of showing progress incrementally.

Envelope example:

{"cap_envelope":1,"capability":"test.focused","outcome":"task-failed","command":"npm test -- --silent","exit_code":1,"signal":null,"duration_ms":8421,"reason":null,"output_tail":"FAIL src/foo.test.ts\n  âś— should do the thing\n"}

Telemetry

Every cap run — all four outcomes included — appends one JSON line to ~/.dossier/caps.jsonl (append-only, mode 0600; disable with dossier config auditLog false), recording capability, outcome, exit_code, duration_ms, reason (why a non-ok outcome happened), signal, cwd, timestamp, and (non-ok outcomes only, #583) output_tail. This mirrors the runs.jsonl dossier telemetry but stays a separate file because a capability execution is not a dossier run.

When cwd is a git work tree the row also records what was verified (#941): git_head (git rev-parse HEAD), git_tree (HEAD^{tree}) and dirty. The tree is probed before and after the run; dirty is true if either probe found changes (tracked, staged, or untracked-not-ignored files — regardless of status.showUntrackedFiles/submodule-ignore config — or files hidden with --assume-unchanged/--skip-worktree, which includes sparse-checkout trees) or if the run moved HEAD or the tree. The probe strips GIT_* env, uses --no-optional-locks, and has a 10 s total budget; if it times out or errors, the row records git_probe: "timeout"|"error" with dirty: true (and a stderr warning). Rows outside a work tree omit all of these. Rows also record git_prefix (git rev-parse --show-prefix: the directory inside the repo the run happened in) and command_hash (sha256 of the manifest command). The dirty probe covers the whole repo even from a subdirectory (ls-files -v -- :/) and pins core.fsmonitor=false/core.fileMode=true. Every row also records args (the words after --) and args_hash (sha256 of the JSON array).

ai-dossier cap last-ok <id> (--tree <40|64-hex sha> | --here) [--prefix <dir>] [-- <args>] prints the latest clean ok row for that capability, tree AND exact args (gate.test -- --only smoke never satisfies a full-gate lookup). Exit codes: 0 match, 1 no match (no output; dirty, probe-failed, failed, different-args and pre-args_hash rows never match), 2 error (bad --tree, unreadable log), 3 auditLog disabled (cannot answer). A torn line in the log is skipped. The match key is capability + tree + args + directory prefix (default: the current directory’s own; --prefix overrides) + command hash, so a pass from cli/ never satisfies a lookup from the root. --tree must come from a CLEAN tree — --here probes the current directory itself (tree + prefix) and exits 1 when it is dirty. The match key deliberately excludes: git-ignored files, toolchain/CLI versions, environment, and nested repos other than via submodule status.

Capability id vocabulary

Reserved vocabulary for cross-repo consistency (ids are a convention, not enforced — but use these when they fit):

IdMeaning
worktree.prepareCreate/warm a git worktree for development
worktree.cleanupClean up / return a worktree
dependencies.installInstall project dependencies (npm/pnpm/uv/…)
test.focusedFast, targeted test suite (batch member gate fast path)
test.fullComplete test suite (the batch gate’s fallback)
gate.batchThe gate a batch pays once before its PR — normally the same CI-parity gate a single PR pays (#777)
lint.runLinter/formatter check (batch member gate fast path)
typecheck.runType checking (tsc / mypy / …)
build.runBuild the project
environment.startStart dev servers / containers
environment.stopStop dev servers / containers
verify.uiHealth-check a running app before a live UI verification pass drives it

verify.ui and the live UI pass

verify.ui is a doctor, not a launcher — environment.start and environment.stop already own the runtime, and cap run buffers a command’s output and re-emits it after the child exits, so a command that never returns would simply time out as automation-broken.

Its consumer today is imboard-ai/git/review-issue’s Visual Conformance agent, which drives the app in a headless browser on issues the plan phase flagged visual_review=true. That agent needs one thing a 0 exit cannot express — whether writing to this app is safe — so the contract is a token on the last stdout line:

verify.ui exits 0 and prints SCRATCH-DB-OK as its last stdout line when the app answers, its data store is a scratch or test instance, and the outbound side-effect sinks its flows can reach (email, SMS, payments, webhooks, third-party APIs) are sandboxed or disabled.

Anything else — absent capability, non-zero exit, any other last line — and the consumer drives no mutating flow at all. A token is required rather than a convention because the alternative is an agent reading the doctor’s source and forming an opinion about it, and judging a script instead of reading a signal is exactly the substitution a live verification pass exists to remove.

How the batch member gate consumes these (#625)

sched’s per-member gate runs typecheck.run then test.focused after every batch member reports review-done. It degrades exactly as the vocabulary above implies — declaring them is an optimization, never a prerequisite:

OutcomeGate behaviour
okthe member passes that half
task-failed with failing-test evidencethe member is evicted (#594)
task-failed with no evidencethe batch BLOCKS — a capability that produced nothing did not earn a failure (#594)
automation-brokenthe batch BLOCKS — declared, but its machinery could not be trusted (#583/#585)
automation-broken whose reason is command timed out after <N>msthe gate declines the member — recorded capability-unavailable and journalled gate-skipped-timeout:<id>; the parent’s expensive stage covers it (#681). A timeout is a statement about duration, not about the harness’s reliability — and a member touching two workspace roots selects a dependents closure that approaches the whole workspace, so the gate was never “focused” there. The resolved package selection (the capability’s own output) is preserved in the per-gate log and member_gates.output_tail, with the capability’s duration_ms beside it
capability-unavailablethe check is skipped and journalled gate-skipped:<id>; the member is judged on whatever else is available (#625)

A repo that declares neither id runs batches end to end. Skipping costs early detection — a bad member’s commit may be built on before anyone notices — but not correctness: the aggregate batch gate (gate.batch, else test.full) still runs before ship, CI still runs on the batch PR, and #562’s attribution still pins a red suite to the member that caused it. Declaring the two ids buys earlier, cheaper failure, which is the whole point of progressive determinism.

The skip is always journalled. A gate that silently does not run is its own trap, and silence must never read as a pass.

The batch gate: gate.batch (#777)

A batch’s validating phase runs ONE aggregate gate over the combined work of every member before the tail agent opens the batch PR. The scheduler resolves it in this order, each tier preferred to the next:

TierWhat runsWhen
0cap run gate.batchthe manifest declares an active gate.batch
1cap run test.fullno active gate.batch (or cap run reported it capability-unavailable)
2dispatch.suite_command from sched configneither capability available
3the repo’s detected test runnernothing above

gate.batch exists because test.full is the wrong gate for a repo whose full suite is too slow to finish: a batch should pay the same gate a normal PR pays, once — e.g. a CI-parity script, affected-scoped over the union of the members’ diffs — not the most expensive suite the repo has. Its command runs with cwd = the batch integration worktree and two extra environment variables:

VariableValue
DOSSIER_BATCH_BASEthe ref the batch branched from, as a fetchable ref — origin/<base_branch> (scope with git diff "$DOSSIER_BATCH_BASE"...HEAD)
DOSSIER_BATCH_IDthe batch id
  gate.batch:
    command: scripts/ci-parity.sh --isolated-db --base "$DOSSIER_BATCH_BASE"
    lifecycle: active
    timeout_ms: 2700000   # its own budget; test.full's is not consulted
    description: CI-parity gate, affected-scoped against the batch base

Its outcomes map to the batch exactly as test.full’s do: ok → green (next member or the tail); task-failed with a parseable vitest JSON report → red, attributed to the offending member; task-failed with no parseable report, or automation-broken (including a timeout) → the batch blocks suite-unreadable, every member commit preserved. A declared capability’s verdict is never replaced by a detected-runner retry.

Attribution needs a report. Member attribution reads a vitest JSON report from the gate’s stdout. A CI-parity script that prints only its own log gives every red run readable: false, so the batch blocks suite-unreadable instead of pinning the failure on a member. To keep attribution, have gate.batch emit a vitest JSON report on stdout (e.g. --reporter=json on its test step). The contract, for gate.batch and test.full alike: stdout carries exactly one JSON document whose first { is { "testResults": [{ "name": "<repo-relative file>", "assertionResults": [{ "status": "failed", "fullName": "…" }] }] } (human output belongs on stderr; paths must be repo-relative, since they are matched against the members’ changed paths). This repo’s own test.full is node scripts/test-report.mjs (#893): it runs each workspace’s vitest and the script tests one at a time with the JSON reporter, merges them into that document, and records a workspace that dies without a report as a failed record on its package.json so the run stays readable. A red run whose report names no member-owned test bisects; one that produces no document at all still blocks suite-unreadable, now with the tail of the failing output in the suite-failed journal detail.

Diff with three dots. origin/<base_branch> moves whenever the batch worktree fetches. git diff "$DOSSIER_BATCH_BASE"...HEAD diffs from the merge-base and is stable; a two-dot diff would pull in upstream changes merged since the batch branched.

Timeout-prone full gates. A repo whose test.full cannot finish inside any reasonable budget should say so with timeout_prone: true. When that is the only full gate — no active gate.batch — sched enqueue refuses to form a new batch and says why, rather than admitting members into a gate that usually ends suite-unreadable. Declare gate.batch, or run those issues as ordinary full cycles. The check reads the manifest in the directory enqueue runs in, and is skipped when --repo names another repository.

Non-goals (per #463)

Automation mining, shadow-compare execution, and generated-automation lifecycle tooling are follow-ups under the Progressive Determinism plan. A shadow entry today is inert: listed by cap list, refused by cap run.


Rendered from docs/reference/capabilities.md in the repository. Edit it there.