AI Dossier

Model Scorecard

Generated: 2026-09-14T05:00:37.968Z | Window: 2026-08-15 → 2026-09-14

Cost, quality, and speed per LLM, joined from runstate trails (GitHub), runs.jsonl (token/cost telemetry), and events.jsonl (dispatch tier, stall/escalation counts). Regenerate with npm run scorecard -- --days 30. See #566.

n is a confidence column, not a metric — a row with n=1 is one data point, not a trend. Read cost/delivered and delivery rate alongside n, never alone.

Per model × repo × tier

ModelRepoTierAgent CLInDeliveredDelivery rateAC metCost/deliveredMedian API-minMedian wall-clock-minReview fixed/issueStallsEscalationsUnverified exits
<unknown>imboard-ai/ai-dossierunknown17635%98% (n=6)N/AN/A89.227.0 (n=6)000
<unknown>imboard-ai/ai-dossiermidclaude900%N/AN/AN/AN/A0.5 (n=8)000
<unknown>imboard-ai/imboard-monorepounknown701826%99% (n=13)N/AN/A124.37.5 (n=20)000
<unknown>imboard-ai/imboard-monorepomechanicalopencode100%N/AN/AN/AN/A0.0 (n=1)000
<unknown>imboard-ai/imboard-monorepomidclaude,opencode400%N/AN/AN/AN/A1.5 (n=4)000
claude-fable-5imboard-ai/imboard-monorepounknown22100%100% (n=2)N/AN/A130.716.5 (n=2)000
claude-fable-5-1imboard-ai/imboard-monorepounknown77100%100% (n=4)N/AN/A182.116.2 (n=6)000
claude-haiku-4-5-20251001imboard-ai/imboard-monorepomechanicalclaude11100%N/A$1.564 (n=1)86.11075.61.0 (n=1)000
claude-opus-5imboard-ai/ai-dossierunknown5480%100% (n=4)N/AN/A52.832.3 (n=4)000
claude-opus-5imboard-ai/ai-dossierstrongclaude22100%100% (n=1)$52.143 (n=2)133.9132.626.0 (n=2)000
claude-opus-5imboard-ai/imboard-monorepounknown383797%99% (n=24)N/AN/A131.819.1 (n=37)000
claude-opus-5imboard-ai/imboard-monorepomechanicalclaude11100%100% (n=1)$25.241 (n=1)396.71616.823.0 (n=1)01 (1.00/issue)1 (1.00/issue)
claude-opus-5imboard-ai/imboard-monorepostrongclaude100%75% (n=1)N/AN/AN/A35.0 (n=1)01 (1.00/issue)1 (1.00/issue)
claude-opus-5[1m]imboard-ai/ai-dossierunknown11100%100% (n=1)N/AN/A45.331.0 (n=1)000
claude-opus-5[1m]imboard-ai/imboard-monorepounknown11100%N/AN/AN/A106.322.0 (n=1)000
claude-sonnet-5imboard-ai/ai-dossierunknown77100%100% (n=5)N/AN/A33.77.3 (n=7)000
claude-sonnet-5imboard-ai/ai-dossiermechanicalclaude11100%100% (n=1)$4.793 (n=1)20.019.72.0 (n=1)000
claude-sonnet-5imboard-ai/ai-dossiermidclaude343191%97% (n=32)$18.341 (n=14)45.349.712.8 (n=34)000
claude-sonnet-5imboard-ai/ai-dossierstrongclaude9889%76% (n=8)$24.682 (n=8)88.082.824.7 (n=9)09 (1.00/issue)6 (0.67/issue)
claude-sonnet-5imboard-ai/imboard-monorepounknown343294%98% (n=30)N/AN/A187.47.8 (n=33)000
claude-sonnet-5imboard-ai/imboard-monorepomechanicalclaude,opencode201995%97% (n=19)$30.576 (n=19)91.0305.814.4 (n=20)022 (1.10/issue)13 (0.65/issue)
claude-sonnet-5imboard-ai/imboard-monorepomidclaude300%100% (n=1)N/AN/AN/A3.0 (n=1)01 (0.33/issue)0
claude-sonnet-5imboard-ai/imboard-monorepostrongclaude500%100% (n=2)N/AN/AN/A9.7 (n=3)05 (1.00/issue)4 (0.80/issue)
deepseek-v4-pro-0813imboard-ai/imboard-monorepounknown33100%86% (n=2)N/AN/A139.90.5 (n=2)000
glm-5.3imboard-ai/ai-dossierunknown161488%100% (n=15)N/AN/A67.115.7 (n=16)000
glm-5.3imboard-ai/imboard-monorepounknown282693%97% (n=24)N/AN/A173.38.7 (n=26)000
glm-5.3imboard-ai/imboard-monorepomechanicalopencode22100%100% (n=2)N/AN/A337.09.5 (n=2)1 (0.50/issue)4 (2.00/issue)0
glm-5.3imboard-ai/imboard-monorepostrongopencode100%N/AN/AN/AN/AN/A02 (2.00/issue)0
glm-5.3-flashimboard-ai/ai-dossierunknown141393%100% (n=9)N/AN/A57.65.0 (n=13)000
glm-5.3-flashimboard-ai/imboard-monorepounknown1717100%95% (n=16)N/AN/A215.29.4 (n=17)000
gpt-5.6-lunaimboard-ai/ai-dossierunknown100%N/AN/AN/AN/AN/A000
gpt-5.6-lunaimboard-ai/imboard-monorepounknown5480%100% (n=3)N/AN/A255.31.5 (n=4)000
gpt-5.6-terraimboard-ai/ai-dossierunknown12758%98% (n=8)N/AN/A32.83.9 (n=8)000
gpt-5.6-terraimboard-ai/imboard-monorepounknown14964%100% (n=2)N/AN/A102.11.7 (n=9)000
gpt-5.6-terraimboard-ai/imboard-monorepomechanicalopencode11100%100% (n=1)N/A78.1886.82.0 (n=1)000
gpt-6-astraimboard-ai/imboard-monorepounknown2150%92% (n=2)N/AN/A55.110.0 (n=2)000
kimi-k3imboard-ai/imboard-monorepounknown100%N/AN/AN/AN/AN/A000
kimi-k3-fastimboard-ai/imboard-monorepounknown9889%97% (n=9)N/AN/A353.74.8 (n=9)000
kimi-latestimboard-ai/imboard-monorepomechanicalopencode11100%100% (n=1)N/AN/A527.115.0 (n=1)000

Totals per model (all repos/tiers)

One row per model, with the gateways it was served through as ↳ sub-rows whenever there is more than one — the fold that makes the model row readable would otherwise hide a gateway costing more or delivering less than the same weights elsewhere.

Billable tokens count uncached input + cache-creation + cache-read + output — cache reads are billed and, on this fleet, are the dominant term (issue #540: 262 uncached vs 13.4M cache-read). That is the same total batch-pilot-2-execution.md §13 publishes.

ModelProviderAgent CLInDeliveredDelivery rateΔ vs prevAC metCost/deliveredBillable tokens/deliveredMedian API-minMedian wall-clock-minReview fixed/issueStallsEscalationsUnverified exits
<unknown>directclaude,opencode1012424%-0pt98% (n=19)N/AN/AN/A114.88.3 (n=39)000
claude-fable-5directunknown22100%+0pt100% (n=2)N/AN/AN/A130.716.5 (n=2)000
claude-fable-5-1directunknown77100%+100pt100% (n=4)N/AN/AN/A182.116.2 (n=6)000
claude-haiku-4-5-20251001directclaude11100%—N/A$1.564 (n=1)7,602,537 (n=1)86.11075.61.0 (n=1)000
claude-opus-5directclaude474494%-1pt99% (n=31)$43.176 (n=3)60,985,501 (n=3)198.0130.821.0 (n=45)02 (0.04/issue)2 (0.04/issue)
claude-opus-5[1m]directunknown22100%+0pt100% (n=1)N/AN/AN/A75.826.5 (n=2)000
claude-sonnet-5directclaude,opencode1139887%-0pt96% (n=98)$24.761 (n=42)51,348,820 (n=42)56.975.711.9 (n=108)037 (0.33/issue)23 (0.20/issue)
deepseek-v4-pro-0813directunknown33100%+0pt86% (n=2)N/AN/AN/A139.90.5 (n=2)000
glm-5.34 providers ↓opencode474289%+0pt98% (n=41)N/AN/AN/A139.311.3 (n=44)1 (0.02/issue)6 (0.13/issue)0
↳directunknown353291%—98% (n=32)N/AN/AN/A139.711.5 (n=33)000
↳llmgatewayunknown8788%—100% (n=7)N/AN/AN/A103.412.1 (n=8)000
↳z-aiopencode3267%—100% (n=2)N/AN/AN/A337.09.5 (n=2)1 (0.33/issue)6 (2.00/issue)0
↳zai-coding-planunknown11100%—N/AN/AN/AN/A139.32.0 (n=1)000
glm-5.3-flash2 providers ↓unknown313097%+4pt97% (n=25)N/AN/AN/A160.67.5 (n=30)000
↳directunknown2121100%—96% (n=20)N/AN/AN/A158.36.3 (n=21)000
↳zai-coding-planunknown10990%—100% (n=5)N/AN/AN/A173.710.2 (n=9)000
gpt-5.6-luna2 providers ↓unknown6467%-13pt100% (n=3)N/AN/AN/A255.31.5 (n=4)000
↳directunknown3267%—100% (n=2)N/AN/AN/A364.61.0 (n=2)000
↳openaiunknown3267%—100% (n=1)N/AN/AN/A165.02.0 (n=2)000
gpt-5.6-terra2 providers ↓opencode271763%-12pt98% (n=11)N/A23,354,866 (n=1)78.169.02.7 (n=18)000
↳llmgatewayunknown3267%—N/AN/AN/AN/A119.30.5 (n=2)000
↳openaiopencode241563%—98% (n=11)N/A23,354,866 (n=1)78.169.02.9 (n=16)000
gpt-6-astraopenaiunknown2150%—92% (n=2)N/AN/AN/A55.110.0 (n=2)000
kimi-k3directunknown100%+0ptN/AN/AN/AN/AN/AN/A000
kimi-k3-fastdirectunknown9889%+0pt97% (n=9)N/AN/AN/A353.74.8 (n=9)000
kimi-latestopenrouteropencode11100%+0pt100% (n=1)N/AN/AN/A527.115.0 (n=1)000
TOTAL—claude,opencode40028471%—97% (n=249)$25.458 (n=46)50,437,540 (n=47)64.0113.211.5 (n=313)1 (0.00/issue)45 (0.11/issue)25 (0.06/issue)

Of the 46 delivered issues with a cost figure, 17 were recovered from the dispatch’s own agent log because runs.jsonl recorded none — see Limitations.

Wall-clock per phase (all models)

Median seconds between a phase’s milestone and the previous one, from the trails’ own at= stamps. Wall-clock is the only per-phase measure available: cost AND API-minutes both come from one agent session that usually spans several phases, so neither can be attributed to a phase (see Limitations).

PhasenMedian
batch-review650.2m
batch-validate1735.6m
gate8124.9m
implement31021.8m
merge-wait21520.0m
review30820.0m
batch-report57.5m
plan3135.7m
batch-ship53.0m
ship2932.8m
setup3102.3m
report25758s
classify111s

Reconciliation

This is a regenerated snapshot, not the first one — see git history for docs/reports/model-scorecard.md for prior windows. The first-snapshot reconciliation against batch-pilot-2-execution.md §13.3 and model-agnostic-fleet.md ran once, at #566.

Limitations

  • A moving version tag folds onto its pin only where someone declared the mapping. glm-latest → glm-5.3 is declared (MODEL_ALIASES in cli/src/runstate-stats.ts, the mapping #566 states) and folds, along with every routed spelling of it. Which pin a -latest tag points at is a fact about the provider’s state, not about the string, so it cannot be derived here and it goes stale the moment the provider ships a new version under the same tag — when that happens, update the value in MODEL_ALIASES rather than reading the row as one version. Undeclared tags (kimi-latest, which has both kimi-k3 and kimi-k3-fast as plausible pins) keep their own row and are named in Data warnings — a guessed alias misattributes cost and quality silently, a missing one only splits a row.
  • A context-window variant keeps its own row. claude-opus-5[1m] does not fold into claude-opus-5: the milestone protocol says the suffix should never have been written (gate records the bare model id), but 1M-context is billed differently, so folding it would blend two cost profiles to fix a formatting slip. Read the two rows together when judging quality, separately when judging cost.
  • Cost comes from two sources, and the column says which. ~/.dossier/runs.jsonl is authoritative; where a dispatch predates the telemetry fix (#564) and left it null, the figure is recovered from that dispatch’s own agent log under ~/.dossier/sched/<slug>/runs/. Both are on-host only — a dispatch run from another machine has neither, and its row reads N/A because the data is elsewhere, not because it was free.
  • Δ vs prev is snapshot-over-snapshot, not a strict 7-day delta. It compares this run against whatever sidecar is on the base branch — two overlapping 30-day windows when the weekly cron produced both, and a longer gap whenever a weekly PR went unmerged. A baseline older than the window raises its own data warning naming the age, so the column is never silently stale; read it as “since the last published snapshot”.
  • The window selects ISSUES by last update, not runs by date. The trail read is gh issue list --search "updated:>=<start>", and a matched issue contributes its whole dispatch history — so an old run on an issue merely touched inside the window counts in full. Read the window as “every run on an issue active in the last N days”, not “every run started in the last N days”.
  • A <unknown> tier is usually an artifact of that mismatch, not an untiered dispatch. Tier and agent CLI come from events.jsonl events inside the window, while the trail data behind the same row is not time-filtered — so a run whose dispatch event predates the window keeps its outcome columns and loses its tier. The per-model totals fold tiers away and are unaffected; only the per-tier rows split.
  • Cost and API-minutes per phase are not separable. A dispatch is usually one continuous agent session covering several phases, so both are recorded per issue, not per phase — the per-phase section above reports wall-clock only, which the milestone at= stamps do carry. For a per-phase breakdown of a single run rather than a median across many, run ai-dossier runstate stats directly.
  • Stall/escalation/unverified-exit counts are a per-host gap. They come from ~/.dossier/sched/<project>/events.jsonl, which only exists on the machine that ran the dispatch. A run dispatched from another host reports 0 for these columns here even if it really stalled — fleet-cli-audit.sh documents which hosts exist; this script does not collect across them.
  • Cost/delivered averages only delivered issues. Work that was dispatched and then blocked, evicted, or abandoned is not in the denominator or the numerator — see the n vs Delivered columns for how much of a bucket that excludes.

Data warnings

  • imboard-ai/ai-dossier: 33 run(s) recorded no model= — their row is bucketed as , and its outcome columns are not attributable to any model
  • imboard-ai/ai-dossier: 26 of 166 run(s) are classifier dispatches (classify milestones only) — excluded from the by-model and by-class tables, since they record no model= and never ship
  • imboard-ai/ai-dossier: 111 of 140 run(s) (79%) carry no classifier risk= verdict — they are bucketed as , so the class rows describe only the classified remainder
  • imboard-ai/imboard-monorepo: issue imboard-ai/imboard-monorepo#3684 run r-3684-a0cf: gate milestone has an unusable at= value (’$(date’) — skipped, and the next phase’s duration is reported as unknown
  • imboard-ai/imboard-monorepo: 51 run(s) recorded no model= — their row is bucketed as , and its outcome columns are not attributable to any model
  • imboard-ai/imboard-monorepo: ‘kimi-latest’ is a moving version tag with no declared pin — its 1 run(s) sit in their own row, apart from whatever pinned version the tag resolves to; add it to MODEL_ALIASES to fold them
  • imboard-ai/imboard-monorepo: 71 of 320 run(s) are classifier dispatches (classify milestones only) — excluded from the by-model and by-class tables, since they record no model= and never ship
  • imboard-ai/imboard-monorepo: 203 of 249 run(s) (82%) carry no classifier risk= verdict — they are bucketed as , so the class rows describe only the classified remainder

Rendered from docs/reports/model-scorecard.md in the repository. Edit it there.