Debugging and presentation for AI apps

Watch an AI app work, step by step, right beside it.

Pick a scenario and replay a real run. Beside the app, the bench lights up every step it takes, the exact prompt the model was given, and what came back.

Not technical? Enter the lab and start the Slack Helpdesk example: watch what the AI does with a request, step by step, in plain words.

2 scenarios·every prompt, in full·works with OpenTelemetry
In the lab

Pick a scenario to start

Shipped products and an example scenario, all running from recordings of real model calls. Start one, use the app, and watch the bench beside it. Open the lab →

On the bench

What you'll see

Two views of the same run. Presentation tells it in plain words, for anyone in the room. Engineering shows every number underneath, and every number is real.

◆

What it did

Every step the app could take. The path it actually took lights up; the branches it skipped stay dashed, each with a note on why.

Engineering: the graph read from the app's code
⌕

What it looked at

Everything it could look at, what it was actually given, and what its answer rests on. Open any of it, and the exact words the AI was given.

Engineering: system · messages · raw output
✓

How it was checked, who signed off

Each check and whether it passed. When a person must approve, the run waits for them, and says who decided.

Engineering: per-step latency waterfall
$

Time and cost

How long the AI worked, how long it waited for a person, and what it cost, next to how long the job takes by hand.

Engineering: tokens · cost, actual vs estimated
Under the hood

How it works

An app reports what it does as it runs, and each run carries a map of every step it could take, read from the app's own code. The bench draws the map, then fills it in live as the run arrives, or replays a recording with its original timing.

It sits on top of the tracing an app already has, OpenTelemetry or its own events, so it works with whatever you use to observe AI in production. It never changes the app; approvals and every other decision stay where they belong, in the app.

The map comes from the app's own graph (LangGraph today, or an Open Agent Spec flow file), so it can't drift from the code. Plain words sit next to the code they describe and are checked against it, and an app can bring its own story: custom panels that explain its runs, the way comments explain code.

The appemits OpenTelemetry spans
The mapevery step and possible branch, from its code
The storyoptional custom panels
The benchflow · prompts · timeline · cost