Agent Fleet Baseline · two weeks · fixed scope

Know which agent workflows deliver before you scale them.

A delivery-assurance baseline for teams running Claude Code, Codex, Cursor, scheduled agents, or large child-agent fleets. Start locally, keep raw transcripts private, and finish with one operating decision.

LOCAL FIRSTThe free audit reads existing transcripts on the machine. Raw prompts do not need to be uploaded.
PERSON-BLINDFleet findings describe runs and agent families, not individual developer performance.
EVIDENCE-BOUNDEDUnknown and unavailable outcomes remain visible instead of being counted as delivery.
Use it when the workflow matters

From adoption metrics to an operating decision.

Seat counts and token dashboards can show that agents are being used. The baseline examines whether the workflows you depend on leave observable delivery evidence, where coverage ends, and which intervention deserves a before/after check.

01 · DELIVERY

Recurring jobs

Separate a completed transcript from successful write evidence and a local target that is present now.

02 · SCALE

Child-agent fleets

Make hidden fan-out visible and reconcile parent-run conclusions with child-run cost and terminal states.

03 · GOVERNANCE

Rollout approval

Give engineering, security, and employee representatives an explicit data boundary without ranking people.

What happens

One workflow. Four bounded steps.

The engagement stays narrow enough to finish and broad enough to reveal whether the problem is delivery, configuration, cost shape, or missing evidence.

DAYS 1-2

Scope

Select machines, harnesses, one workflow, the decision, and what may leave the machine.

DAYS 3-6

Baseline

Run the local analyses, state coverage, and review scheduled and child-run evidence.

DAYS 7-10

Intervene

Choose one workflow, configuration, model, context, or delivery-probe change.

DAYS 11-14

Decide

Review before/after evidence when available and make one explicit operating recommendation.

Evidence boundary

Clear about what the current product can prove.

Included now

  • Claude Code, Codex, and Cursor local transcript coverage
  • Scheduled, workflow, and child-agent run ledger
  • Terminal and delivery-state classification
  • Successful Write, Edit, and NotebookEdit target checks on the local filesystem
  • Token-cost decomposition, including cache-read and child-run volume
  • Configuration and privacy-boundary review

Not claimed

  • Commit or pull-request attribution to a run
  • Deployment or external-action verification
  • Arbitrary promised report completion
  • Reliable probing inside every container or remote workspace
  • Semantic quality or correctness of an artifact
  • Individual developer productivity
PDF

Use the six-page decision guide internally.

Market, positioning, qualification, offer structure, launch sequence, and a practical checklist in one reviewed document.

Download the decision guide
Design-partner offer

Start with one workflow worth fixing.

The first design partners receive a fixed-scope two-week baseline. The fee can be credited against a later annual assurance agreement.

€1,500-€3,000fixed after confirming machines, harnesses, and workflows
  • Coverage and evidence-boundary statement
  • Delivery ledger and cost-shape review
  • Three evidence-backed failure or waste patterns
  • One prioritized intervention
  • Findings session and written recommendation

Prefer email? Write directly.

Request a fit check

Nothing is submitted to this site. Continue opens a drafted message in your email client.

No card is collected. The fit check is 30 minutes.

Questions

Before you involve another stakeholder.

Does session-viz upload agent transcripts?

The first audit runs locally. Raw transcripts and prompt content do not need to leave the machine. Shared fleet findings use a closed, bounded schema.

Is this an employee-productivity assessment?

No. The baseline examines workflows, runs, coverage, delivery evidence, and cost shape. It does not rank developers or create a person column in fleet telemetry.

What must the customer provide?

A technical owner, one workflow worth examining, access to the relevant local transcripts, and a decision the evidence could change.

What happens if the local audit finds no decision-quality signal?

There is no paid baseline proposal. The audit is the qualification step, not a commitment device.