Delivery assurance for coding agents

Your coding agents said done. Did anything ship?

session-viz verifies Claude Code, Codex and Cursor runs against the files, commits and outputs they were meant to produce. Catch silent delivery failures, see what autonomous work cost, and improve the agent definition — without creating a leaderboard for developers.

Free local audit · MIT licensed · no account required · no person column

Delivery audit · local reportHow 60 days of coding-agent runs actually ended
1,166 runs60 days1 machine
completed · structured completed · prose not resolved — shown as a hole, never guessed truncated zombie Measured proportions, drawn as 680 cells — not one cell per run

What the delivery audit found on one machine

0 / 20successful writes

A scheduled job ran for 30 days, cost real money, and recorded no successful write result. It announced the same next instalment on every run. Every conventional signal was green — coherent transcripts, correct plans, no errors on five of the runs — and a human walked back into twelve of those sessions without noticing.

95.7%of tokens

Almost your entire bill is cache-read — context replayed to the model — not text it generated. It never appears in a per-session view, and on this machine it was 24.19B tokens against 85M of output.

1,102invisible runs

Subagents are the fleet nobody watches. They were 94% of all runs and 99% of autonomous spend. One parent session spawned 347 children; the median parent spawns 26.

8.0% vs 1.2%schema failure

One agent definition failed its caller's schema seven times more often than another. The 95% intervals do not overlap, so the difference is real — and it is a property of a file you can edit, not of a person.

Two planes that cannot join

Fleet observability and a private place to learn are only compatible if the boundary between them is a property of the schema rather than a promise in a policy document.

Schematic. The agent families and their proportions on plane A are from the reference corpus; the three-person team on plane B illustrates the supported topology, not a measured headcount.

Plane A — telemetry

Person-blind by construction

  • Keyed on repo, task, agent definition, config, version — never on a person
  • Closed contribution schema: enums and bounded integers, no free-text field exists
  • k-gated cells; suppressions are counted, never silent
  • Signed reference table served with no request body and no cookie
no shared key
no foreign key to telemetry
boot fails if one appears
Plane B — collaboration

Identity-bearing, on purpose

  • Federated Obsidian vaults: [[wikilinks]] resolve across projects and people
  • Task handoff with an explicit state machine — nothing lands on you unaccepted
  • Live sync over SSE, and a stateless MCP server for agents
  • Note bodies stay on your machine; only titles, tags and link names are indexed

You cannot hand a task from one person to another without knowing both. So the collaboration plane knows them — and is structurally unable to be turned back into per-person performance measurement, because the join does not exist: the telemetry table is keyed on task classes and version bands, the collaboration tables on actors and vaults, and there is no column in common to join on. A boot-time check reads the live schema and refuses to start if any table gains a foreign key to the telemetry table, or if that table gains a column whose name looks like a person's.

What it refuses to tell you

Any per-person rate

Structurally uncomputable. There is no person column to group by, so this is not a setting that can be turned on later.

That the newer model is better

Adoption replaced the old model wholesale, so there is no week where both ran at volume. The difference is unattributable — not weakly, but not at all.

That your prompting improved

On 1,075 turns, not one prompt-form signal survived a workload control. The raw correlations were inverted by task difficulty; naming a file appeared to double friction.

A trend across a format change

The transcript format moved twelve times in sixty days. Any window straddling a change point is returned blocked rather than rendered.

How session-viz verifies the result

Four stages. Everything before the last one runs on your machine, and the last one is optional.

01

Read

Claude Code already writes a JSONL transcript for every session, cron run and subagent. The parser streams all of them — about 600 MB in two seconds — so there is no daemon, no cache to go stale and nothing to install beyond the plugin.

  • Human sessions, scheduled runs, subagents and workflow agents become one run record
  • Worktrees fold back into their repository
  • Secrets are redacted from prompt text before anything is written
02

Classify

Each run gets a terminal state and a delivery state from what the transcript actually witnessed — not from a model's opinion of it. A run that ends on a StructuredOutput call is a success; a naive rule would call it a failure and condemn most of a healthy fleet.

  • Seven terminal states, and unknown stays unknown
  • wrote_ok stays a transcript fact; a separate probe reports whether its target exists locally now
  • Permission denials are separated from model failures
03

Gate

Every comparison passes a statistical gate before it is shown. Signals are stratified by workload, models are only compared within weeks both ran at volume, and trends spanning a transcript-format change are blocked rather than drawn.

  • Two-proportion z-test with an incident floor, not just a sample-size floor
  • k-gated cells; at fleet scale a cell also needs several distinct tenants
  • Whatever fails the gate is listed with the reason, in the Null Book
04

Act

Findings become tickets, tickets become commits. The optional cloud plane adds federated vaults, task handoff between people, live sync and a stateless MCP endpoint so your agents can use all of it directly.

  • Config changes are dated events, which makes before/after measurable
  • [[wikilinks]] resolve across projects and across people
  • Handoffs need explicit acceptance — nothing lands on anyone unasked
Run wall

Every run as one cell, painted only by how it ended. Nothing is aggregated, so nothing can be inflated — and the runs the classifier cannot name are drawn as visible holes rather than smoothed away.

Delivery ledger

What each run reported, plus whether successful Write, Edit and NotebookEdit targets exist on the local filesystem now. Missing local evidence stays distinct from failed delivery; git commits and arbitrary promised outputs are not independently verified yet.

Cost decomposition

Four token classes, shown linear and log side by side because either alone misleads. Cache-read is almost the whole bill and appears in no per-session view.

Knowledge graph

What connects your repositories, built from packages imported and CLIs run rather than from prompt wording. Topics that span everything are dropped: they describe your toolchain, not a relationship.

Substitution canary

A detector whose only job is to tell you an improvement was the instrument going blind. It has already caught one in the reference corpus.

Team and tenancy

Customer workspaces with their own admins, members they invite, and passwordless sign-in by email code or passkey. No view anywhere resolves to a person's performance.

Start with a local audit. Add a fleet when it earns its keep.

The parser, delivery checks and statistical gates stay open and local. Pay when you need a managed team workspace, a repeatable fleet baseline, or somebody accountable for keeping the extraction current as agent formats change.

Local audit

Verify recent runs on your machine before buying anything.

FreeMIT · self-hosted
  • All six local commands: /qcost /qtrends /qruns /qpact /qship /qdoctor
  • The control center and every gate
  • Backend deployable to your own Railway, Fly or box
  • Reference table works offline forever
  • You operate it, you upgrade it
  • No SLA, best-effort issues
Audit recent runs

No local analysis is held back. A user who never contributes anything still gets the full local report, and /qbl is local unless you ask for --shared. The other five commands — /qcontrib, /qsetup, /qfeed, /qshare, /qteam — talk to a workspace, so on RYO they talk to yours.

Hosted

Fleet

A managed workspace for teams operating agents at volume.

Per agent-runusage-based · not per seat
  • Managed backend, Postgres and backups
  • Federated vaults across projects and people
  • Task handoff, live sync, hosted MCP endpoint
  • Signed reference table with a stable key
  • EU region, DPA, deletion on request
  • Still no person column — that is not a tier
Explore the fleet baseline Request beta access

Hosted workspaces are currently a private beta. Ask about access to the collaboration plane, hosted MCP, and signed reference table. No card is collected.

Priced on agent-runs because that is the unit the product measures. Counting people would require identifying them, which the telemetry plane is built not to do.

Assurance

For teams that need a rollout another stakeholder can approve.

Annualscoped engagement
  • Named contact, response-time commitment
  • Deployment review and upgrade assistance
  • Extraction contract pinned to your CLI versions
  • Co-determination pack: what is measured, what cannot be
  • Custom detectors for your recurring tasks
  • Priority on parser fixes when the format moves
Discuss assurance

The transcript format changed twelve times in sixty days. Keeping a parser current is the single most valuable thing to outsource.

Audit your recent agent runs

Install where you work, restart the harness, then run /qruns for the fastest delivery audit. It reads every harness below wherever you install it — the transcripts are on the same disk either way — so a Claude Code install still analyses your Codex and Cursor history. Installing is the per-harness part.

claude plugin marketplace add QSchlegel/session-viz
claude plugin install session-viz@session-viz

Two steps on purpose: install cannot resolve a plugin from a marketplace this machine has never added.

# one session, with a tuned /compact line
$ /qpact

# every session on the machine
$ /qtrends

Runs entirely locally. 600 MB of transcript parses in about two seconds, so there is no cache and no daemon. Restart the session first — skills register at startup.

# optional: the collaboration plane
$ curl -sX POST $HOST/v1/mcp \
    -H "authorization: Bearer $TOKEN" \
    -H "x-actor: you" \
    -d '{"jsonrpc":"2.0","id":1,
         "method":"tools/list"}'

vault_list      vault_register
vault_resolve   vault_dangling
task_create     task_offer
task_accept     task_done
task_list       events_recent

Stateless MCP over HTTP JSON-RPC. No session to keep, so it survives being moved between machines mid-request.

# read from every harness on the disk,
# whichever one you installed into

claude-code  ~/.claude/projects
codex        ~/.codex/sessions
cursor       globalStorage/state.vscdb
cloud        not stored locally

A harness that is present but unreadable, and one that is simply absent, are different sentences with different fixes — so the reports distinguish them rather than showing one silence.

Sign in

Passwordless. A code by email, or a passkey if this device has one.

A new address starts its own workspace. Colleagues join it by invitation, so nobody is pooled with you for sharing an email domain.

Self-hosted installs print the code to the server log until an email provider is configured — see the README of session-viz-cloud.