flow-1 is our new model for understanding agent traces, trained with reinforcement learning to identify hard-to-spot failures and explain why they happened. On our benchmark, it matches GPT-6-sol in detection quality while being 23x cheaper.
Teams are shipping agents into increasingly complex workflows and need to understand where those agents succeed or fail. Agent traces now contain hundreds of model calls and tool results, making it hard to debug an agent or even tell whether it went wrong in the first place. We built flow-1 to make deep, high-quality trace investigations affordable at scale.
flow-1 runs inside what we call the Signals agent. It works like a coding agent, treating the trace as a repo and each span as a file. It starts with a preview of the full trace and a user prompt, which can be as broad as “identify any logical errors.” From there, it can grep across spans and read them in full to investigate. When it finds an issue, it returns structured JSON that follows the user’s schema.
| Model | Detection F1 | Description F1 | Precision | Recall |
|---|---|---|---|---|
| flow-1 | .835 | .741 | .817 | .853 |
| Claude Opus 5 (high) | .890 | .848 | .847 | .938 |
| Claude Sonnet 5 (high) | .835 | .773 | .826 | .844 |
| GPT-6-sol (high) | .816 | .728 | .703 | .974 |
| GPT-6-luna (high) | .795 | .638 | .727 | .878 |
We compared flow-1 with other frontier models using the same Signals agent on our benchmark of 523 difficult agent traces across coding, legal, customer support, CRM and more. flow-1 surpasses GPT-6-sol with high reasoning in failure detection quality. Sol flags more traces, making it noisier, while flow-1 digs into the evidence to surface the failures that matter.
The cost of trace analysis includes the full Signals agent run and all its retrieval steps. Flow-1 is priced at $0.05 per million input tokens, $0.01 for cached input and $0.30 for output. To compare costs, we use statistics from Signals agent runs on traces with under 100k total LLM tokens. flow-1 can process 23x as many traces per dollar with sol-level analysis quality. It’s also 25% cheaper than luna, with significantly better failure detection and explanations. Another RL checkpoint is already training to make flow-1 a more efficient thinker and bring costs down further.
| Model | Traces per $1 | Cost per trace |
|---|---|---|
| flow-1 | 888 | $0.0011 |
| Claude Opus 5 | 7 | $0.149 |
| Claude Sonnet 5 | 12 | $0.086 |
| GPT-6-sol | 38 | $0.026 |
| GPT-6-luna | 668 | $0.0015 |
We trained flow-1 on synthetic data derived from realistic agent workflows in industries where failures have real consequences. Here's the composition of the training datasets.
| Industry / workflow | Approx. share of training examples |
|---|---|
| Software engineering | 48.2% |
| Customer support and telecom | 20.3% |
| General web and assistant tasks | 9.2% |
| Legal services | 7.7% |
| CRM and sales operations | 6.5% |
| Enterprise analytics | 3.5% |
| Finance and data analysis | 2.4% |
| Professional knowledge work | 1.3% |
| Healthcare administration | 1.0% |
Flow-1’s core advantage is that it was trained inside the Signals agent itself. Through RL, it learned to use the agent’s tools to investigate traces accurately and efficiently, from start to finish. Let's now look at flow-1 in action.
In one run, a legal agent prepared a contract redline and an issues list. The parties had agreed to a 105-day film availability window. The agent marked the point closed, but left the contract schedule at 90 days. All three models caught the contract error, and flow-1 also tied it to the misleading completion status in the issues list.
flow-1: “The agent never edited the schedule table to the agreed 105-day window, even though the issues list reported it as conformed clean.”
GPT-6-sol: “The redline failed to incorporate the agreed 105-day film availability window.”
GPT-6-luna: “The counter-markup leaves core availability-window provisions incorrect or incomplete instead of implementing the instructed negotiated terms.”
In another trace, a coding agent reported its work complete, even though not all tests have passed. It blamed the remaining test failures on a pre-existing issue. But the logs showed that its explanation covered only one of the failures. flow-1 and sol flagged the misleading test report, while luna didn't identify anything:
flow-1: “only the telemetry test failure was directly linked to the missing _resolve_registry_uri function.”
GPT-6-sol: “The agent incorrectly attributed an end-to-end test failure to a different missing function than the traceback identified.”
GPT-6-luna: No error found.
A code-review agent spotted a change that replaced yaml with pyyaml. It was an actual bug, correctly identified by the review agent. flow-1 and sol correctly found no error in the trace, while luna flagged the finding as unsupported.
flow-1: No error found.
GPT-6-sol: No error found.
GPT-6-luna: “The agent reported a code-review regression based on distribution-to-import-name mappings that its inspected evidence does not establish.”
There are still gaps. In another contract review, flow-1 cleared a redline that left out mandatory AI-governance protections, even though the agent’s own memo said they were required. Both GPT-6 models caught the omission:
GPT-6-sol: “The tracked redline left mandatory AI-governance protections unaddressed despite identifying them as critical.”
Failure detection is the most common use case, but Signals is a general trace processor. You can use it to identify user frustration, extract structured data from conversations and tool results, or understand which requests agents hand off and why. You can simply define what to look for and the schema for the result.
In our Signals benchmark, one signal definition specifically asked for contradictions within deliverables. A legal agent produced a contract redline and an accompanying memo, but their milestone totals differed by $20M. flow-1 and sol caught the mismatch in the totals. Luna found other inconsistencies, but missed the most crucial one:
flow-1: “The issues memo states the milestone total is $190M and the aggregate deal value is $457M, but the redline's milestone ladder sums to $210M, making the aggregate $477M.”
GPT-6-sol: “The redline's revised milestone payments exceed an unchanged payment cap and contradict the memo's aggregate deal figure.”
GPT-6-luna: “The issues memo recommends a first-commercial-sale milestone, but the redline's replacement milestone list omits that milestone.”
As agents take on more work, reading their traces one by one stops being practical. We built flow-1 to make that understanding affordable at scale: where agents fail, what users need and what happens inside complex workflows.