Pick a scenario and replay a real run. Beside the app, the bench lights up every step it takes, the exact prompt the model was given, and what came back.
Not technical? Enter the lab and start the Slack Helpdesk example: watch what the AI does with a request, step by step, in plain words.
Shipped products and an example scenario, all running from recordings of real model calls. Start one, use the app, and watch the bench beside it. Open the lab →
Two views of the same run. Presentation tells it in plain words, for anyone in the room. Engineering shows every number underneath, and every number is real.
Every step the app could take. The path it actually took lights up; the branches it skipped stay dashed, each with a note on why.
Everything it could look at, what it was actually given, and what its answer rests on. Open any of it, and the exact words the AI was given.
Each check and whether it passed. When a person must approve, the run waits for them, and says who decided.
How long the AI worked, how long it waited for a person, and what it cost, next to how long the job takes by hand.
An app reports what it does as it runs, and each run carries a map of every step it could take, read from the app's own code. The bench draws the map, then fills it in live as the run arrives, or replays a recording with its original timing.
It sits on top of the tracing an app already has, OpenTelemetry or its own events, so it works with whatever you use to observe AI in production. It never changes the app; approvals and every other decision stay where they belong, in the app.
The map comes from the app's own graph (LangGraph today, or an Open Agent Spec flow file), so it can't drift from the code. Plain words sit next to the code they describe and are checked against it, and an app can bring its own story: custom panels that explain its runs, the way comments explain code.