Delivery assurance for coding agents
session-viz verifies Claude Code, Codex and Cursor runs against the files, commits and outputs they were meant to produce. Catch silent delivery failures, see what autonomous work cost, and improve the agent definition — without creating a leaderboard for developers.
Free local audit · MIT licensed · no account required · no person column
A scheduled job ran for 30 days, cost real money, and recorded no successful write result. It announced the same next instalment on every run. Every conventional signal was green — coherent transcripts, correct plans, no errors on five of the runs — and a human walked back into twelve of those sessions without noticing.
Almost your entire bill is cache-read — context replayed to the model — not text it generated. It never appears in a per-session view, and on this machine it was 24.19B tokens against 85M of output.
Subagents are the fleet nobody watches. They were 94% of all runs and 99% of autonomous spend. One parent session spawned 347 children; the median parent spawns 26.
One agent definition failed its caller's schema seven times more often than another. The 95% intervals do not overlap, so the difference is real — and it is a property of a file you can edit, not of a person.
Fleet observability and a private place to learn are only compatible if the boundary between them is a property of the schema rather than a promise in a policy document.
Schematic. The agent families and their proportions on plane A are from the reference corpus; the three-person team on plane B illustrates the supported topology, not a measured headcount.
You cannot hand a task from one person to another without knowing both. So the collaboration plane knows them — and is structurally unable to be turned back into per-person performance measurement, because the join does not exist: the telemetry table is keyed on task classes and version bands, the collaboration tables on actors and vaults, and there is no column in common to join on. A boot-time check reads the live schema and refuses to start if any table gains a foreign key to the telemetry table, or if that table gains a column whose name looks like a person's.
Structurally uncomputable. There is no person column to group by, so this is not a setting that can be turned on later.
Adoption replaced the old model wholesale, so there is no week where both ran at volume. The difference is unattributable — not weakly, but not at all.
On 1,075 turns, not one prompt-form signal survived a workload control. The raw correlations were inverted by task difficulty; naming a file appeared to double friction.
The transcript format moved twelve times in sixty days. Any window straddling a change point is returned blocked rather than rendered.
Four stages. Everything before the last one runs on your machine, and the last one is optional.
Claude Code already writes a JSONL transcript for every session, cron run and subagent. The parser streams all of them — about 600 MB in two seconds — so there is no daemon, no cache to go stale and nothing to install beyond the plugin.
Each run gets a terminal state and a delivery state from what the transcript actually witnessed — not from a model's opinion of it. A run that ends on a StructuredOutput call is a success; a naive rule would call it a failure and condemn most of a healthy fleet.
Every comparison passes a statistical gate before it is shown. Signals are stratified by workload, models are only compared within weeks both ran at volume, and trends spanning a transcript-format change are blocked rather than drawn.
Findings become tickets, tickets become commits. The optional cloud plane adds federated vaults, task handoff between people, live sync and a stateless MCP endpoint so your agents can use all of it directly.
Every run as one cell, painted only by how it ended. Nothing is aggregated, so nothing can be inflated — and the runs the classifier cannot name are drawn as visible holes rather than smoothed away.
What each run reported, plus whether successful Write, Edit and NotebookEdit targets exist on the local filesystem now. Missing local evidence stays distinct from failed delivery; git commits and arbitrary promised outputs are not independently verified yet.
Four token classes, shown linear and log side by side because either alone misleads. Cache-read is almost the whole bill and appears in no per-session view.
What connects your repositories, built from packages imported and CLIs run rather than from prompt wording. Topics that span everything are dropped: they describe your toolchain, not a relationship.
A detector whose only job is to tell you an improvement was the instrument going blind. It has already caught one in the reference corpus.
Customer workspaces with their own admins, members they invite, and passwordless sign-in by email code or passkey. No view anywhere resolves to a person's performance.
The parser, delivery checks and statistical gates stay open and local. Pay when you need a managed team workspace, a repeatable fleet baseline, or somebody accountable for keeping the extraction current as agent formats change.
Verify recent runs on your machine before buying anything.
No local analysis is held back. A user who never contributes anything still gets the full local report, and /qbl is local unless you ask for --shared. The other five commands — /qcontrib, /qsetup, /qfeed, /qshare, /qteam — talk to a workspace, so on RYO they talk to yours.
A managed workspace for teams operating agents at volume.
Hosted workspaces are currently a private beta. Ask about access to the collaboration plane, hosted MCP, and signed reference table. No card is collected.
Priced on agent-runs because that is the unit the product measures. Counting people would require identifying them, which the telemetry plane is built not to do.
For teams that need a rollout another stakeholder can approve.
The transcript format changed twelve times in sixty days. Keeping a parser current is the single most valuable thing to outsource.
Install where you work, restart the harness, then run /qruns for the fastest delivery audit. It reads every harness below wherever you install it — the transcripts are on the same disk either way — so a Claude Code install still analyses your Codex and Cursor history. Installing is the per-harness part.
claude plugin marketplace add QSchlegel/session-viz claude plugin install session-viz@session-viz
Two steps on purpose: install cannot resolve a plugin from a marketplace this machine has never added.
# one session, with a tuned /compact line $ /qpact # every session on the machine $ /qtrends
Runs entirely locally. 600 MB of transcript parses in about two seconds, so there is no cache and no daemon. Restart the session first — skills register at startup.
git clone https://github.com/QSchlegel/session-viz cd session-viz && npm install && npm run build node plugins/session-viz/scripts/install.mjs codex
Codex reads the same skill format Claude Code does. What is not portable is the line inside each skill that runs the analysis — it names ${CLAUDE_PLUGIN_ROOT}, a variable only Claude Code sets, which anywhere else expands to nothing. The installer resolves it on the way in.
# what is installed, and is it current $ install.mjs --list # see it before it writes $ install.mjs --dry-run # removes only what it wrote $ install.mjs --uninstall codex
Skills are copied, scripts are not — copies point back at the checkout, so re-run after updating. Codex has no disable-model-invocation, so it may run these itself rather than only when asked. They are read-only analyses; read-only is not the same as expected.
git clone https://github.com/QSchlegel/session-viz cd session-viz && npm install && npm run build node plugins/session-viz/scripts/install.mjs cursor
With no argument the installer takes every harness it finds on the machine. Naming one is the explicit form.
# Cursor keeps no per-session file and no cwd: # one SQLite database for the whole machine, # and the repository is reconstructed from # the files each conversation touched. 641 conversations here, reaching back nine months earlier than the rest of the corpus.
Cursor records a token count on a minority of messages, so its spend is reported as a floor, not a total — and every report says so rather than showing the gap as a zero.
# nothing to install — and nothing to read $ claude agents --json # lists local runs only
Sessions run on claude.ai/code keep their transcripts server-side. They are not under ~/.claude, not in the desktop app's support directory, and no local file holds them.
# attach from this machine and it writes a # local transcript like any other session $ claude --cloud <session_id> # or point at an export $ export SESSION_VIZ_TRANSCRIPTS=cloud=/path
This cannot be fixed by reading harder, so every report states the absence instead of quietly leaving a surface you use out of the numbers.
# optional: the collaboration plane $ curl -sX POST $HOST/v1/mcp \ -H "authorization: Bearer $TOKEN" \ -H "x-actor: you" \ -d '{"jsonrpc":"2.0","id":1, "method":"tools/list"}' vault_list vault_register vault_resolve vault_dangling task_create task_offer task_accept task_done task_list events_recent
Stateless MCP over HTTP JSON-RPC. No session to keep, so it survives being moved between machines mid-request.
# read from every harness on the disk, # whichever one you installed into claude-code ~/.claude/projects codex ~/.codex/sessions cursor globalStorage/state.vscdb cloud not stored locally
A harness that is present but unreadable, and one that is simply absent, are different sentences with different fixes — so the reports distinguish them rather than showing one silence.