An AI agent evaluation platform runs repeatable tasks, applies code checks, model judges, or human review, and records results across versions. For an agent, a useful evaluation can check the final answer, tool decisions, and the resulting state: whether the right order was refunded, for the right amount, exactly once.
A reply saying “Your refund is complete” does not establish that the payment succeeded. A recorded tool call does not establish it either. Outcome checks establish what changed; a trace helps you investigate how it happened.
This guide compares nine tools and product families on evaluation workflow, trace access, regression analysis, production monitoring, and deployment. Written by the Laminar team, it uses Laminar as the worked example. Capabilities were checked against primary documentation on October 5, 2026. Recommendations reflect workflow fit, not a common benchmark of grading accuracy, speed, or cost.
Scores you can investigate and trust
Trace linkage matters when a low score needs explanation. An executor trace can expose a bad argument, a failed lookup, or a retry. Evaluator records should show what evidence the grader received, its rubric or code version, its model configuration, and its output. An explanation, when available, is useful for review; it is not guaranteed access to a model's internal reasoning.
Visibility and validity are separate. A grader can read a complete trace and still be wrong. A deterministic check of a payment ledger can correctly establish the outcome without seeing the agent's internal steps. Use both where the task requires them.
During a trial, inspect one failed example. Check whether you can reach its execution and grader inputs, distinguish missing instrumentation from missing actions, and verify the score yourself. Recording custom tools, remote services, and subagents may require explicit instrumentation and trace-context propagation.
What to compare in an agent evaluation platform
| Criterion | What to check |
|---|---|
| Outcome and trajectory checks | Can you score tool arguments and authoritative final state? Can you allow multiple valid paths? |
| Grader validity | Can you inspect and version graders, calibrate model judgments, and review disagreements? |
| Regression analysis | Can you compare the same cases across versions and inspect individual changes? |
| Failure to test case | Can production evidence become a reviewed, reproducible task with a correct target? |
| Production coverage | What evidence is analyzed, which traffic is excluded, and are results scores, events, or interventions? |
| Workflow and deployment | Does it fit your SDK, CI, review process, data-residency needs, and budget? |
Simulation adds conversations and scenarios that existing traffic may not cover. It helps before launch and when testing new features, rare failures, or adversarial behavior after launch. A simulated user's behavior is itself an assumption to validate.
Best AI agent evaluation platforms in 2026
Laminar
Laminar evaluations are SDK-first: pass datapoints, an executor, and evaluator functions to evaluate(). Each datapoint has an evaluation trace with executor and evaluator spans. Instrumented LLM and tool calls appear under the executor; return intermediate values when a scorer needs them.
Evaluators can use code or model calls and return one score or several dimensions. The run view places datapoints beside their traces. Related runs share a group for progression charts and per-datapoint comparisons. Keep the dataset and its ordering stable when comparing versions.
A production span can become a dataset entry, with labeling queues supporting review. An observed output needs correction before it serves as a reference. Signals provide a separate production-analysis workflow, described below.
Fit: developers who want custom scoring connected to debugging. Limits: bring your own task-specific evaluators and simulation. The documented evaluation runner does not automatically reuse those functions on live traffic. The core platform is open source; self-hosted Signals, clusters, and alerts require an enterprise license. Deployment comparison.
Braintrust
Braintrust combines datasets, experiments, tracing, and scorers. Autoevals supplies reusable scoring functions; custom scorers let teams express application-specific requirements. Its online scoring supports individual spans, complete traces, or groups of traces, with filters and sampling.
This supports evaluation of intermediate actions and multi-turn interactions, not only final answers. Scorer spans make the scoring inputs and outputs inspectable. Braintrust suits teams seeking a managed experiment and production-scoring workflow. The platform's deployment terms are separate from its publicly available SDKs and Autoevals library.
LangSmith
LangSmith supports dataset experiments, code and model evaluators, pairwise comparisons, and human review. Online evaluation can assess production runs and multi-turn threads. Its evaluation documentation distinguishes when evaluation runs from how a grader is implemented.
LangChain and LangGraph integrations make it a natural shortlist candidate on those stacks. Multi-turn simulation is available through the OpenEvals SDK workflow, and Insights Agent groups production interactions into patterns. These are distinct capabilities: simulation generates interactions, evaluation grades them, and clustering organizes observed behavior. Check your plan for managed and self-hosted features.
Langfuse
Langfuse is an open-source tracing and evaluation platform with cloud and self-hosted options. Its evaluation workflow connects production traces, annotation, datasets, and experiments. Teams can use model judges, human scores, or custom evaluation logic.
Its experiment GitHub Action can run in CI; the script must enforce the intended regression threshold. Langfuse is a useful candidate when self-hosting, production scoring, and collaborative review matter. Review the current deployment and feature terms separately from the availability of the source code.
Arize Phoenix and Arize AX
Phoenix provides OpenTelemetry-based tracing, datasets, experiments, and code or model evaluation. Its evaluation tools include prebuilt checks for tasks such as retrieval and tool use, with evaluator tracing for inspection. Arize AX is a separate enterprise offering, so its monitoring capabilities should not be assumed to ship in Phoenix.
Phoenix is self-hostable and source-available under Elastic License 2.0. That license includes restrictions on providing it as a hosted service; it is not equivalent to an OSI-approved open-source license. Compare the product and license you would actually deploy.
Maxim
Maxim emphasizes personas, scenarios, multi-turn simulation, and evaluation. Its simulation offering combines prebuilt and custom evaluators, run comparisons, human review, and CI integration. Production-derived and synthetic datasets are both part of its workflow.
It is a useful shortlist candidate when conversational coverage is the main concern. Test whether its simulated users reproduce your domain's difficult interactions, then compare findings with human and production evidence.
DeepEval and Confident AI
DeepEval is an open-source framework for tests in Python, with agent metrics for tool correctness, argument correctness, task completion, and plans. Its tool correctness metric has configurable matching criteria: defaults should not be mistaken for checking every argument, output, or side effect.
DeepEval also provides a ConversationSimulator in the framework itself. Confident AI adds a managed platform for collaboration and evaluation workflows. This pairing suits teams who want tests alongside application code; distinguish local framework features from cloud services.
Promptfoo
Promptfoo is an open-source CLI for declarative evaluation and adversarial testing. Its tracing support uses OpenTelemetry, correlates traces with test executions, and exposes them to assertions and grading. A built-in receiver supports development and testing without a separate collector.
It is therefore more than an output-only test runner. Internal execution visibility still depends on the provider and application instrumentation. Shortlist it for repository-based test configurations and multi-turn red teaming; assess ongoing production monitoring separately.
Galileo
Galileo combines tracing, experiments, model-based metrics, and runtime controls. Luna evaluators and custom metrics support quality assessment, while Agent Control evaluates LLM and tool inputs and outputs during execution.
Its older Protect product is deprecated as of June 2026; current documentation recommends Agent Control for new setups. Galileo is a candidate when evaluation must connect to runtime policy enforcement. Measure false positives, missed violations, and latency for your controls; a guardrail does not guarantee prevention.
Comparison: agent evaluation workflows side by side
This table summarizes documented routes, not an exhaustive list of features. “Via SDK” means development work is involved. The paired product rows distinguish a local tool from its commercial companion.
| Tool or product family | Evaluation and inspection | Production route | Simulation route highlighted here |
|---|---|---|---|
| Laminar | Custom functions, datapoint traces, grouped comparisons | Signals, events, clusters, alerts | Bring your own |
| Braintrust | Experiments, scorers, traces and scorer spans | Span, trace, and grouped online scoring | Not assessed here |
| LangSmith | Experiments, trajectory checks, human review | Online evaluators and Insights Agent | OpenEvals SDK |
| Langfuse | Dataset experiments, trace scores, annotation | Production scoring and analytics | Not assessed here |
| Phoenix / AX | Phoenix traces, datasets, experiments | Phoenix can score collected data; assess AX monitoring separately | Not assessed here |
| Maxim | Simulation runs, evaluators, comparisons | Online evaluation | Personas and multi-turn scenarios |
| DeepEval / Confident AI | Framework metrics and traces; managed collaboration | Confident AI platform | DeepEval ConversationSimulator |
| Promptfoo | Assertions, test matrix, OpenTelemetry traces | Assess separately from its test runner | Adversarial multi-turn testing |
| Galileo | Experiments, trace metrics | Production metrics and Agent Control | Not assessed here |
For deployment, Laminar and Langfuse offer self-hosted platforms; Phoenix is self-hostable under ELv2; DeepEval and Promptfoo can run locally. A local framework is not the same as a self-hosted collaboration service. Verify enterprise entitlements, retention, and residency requirements for the configuration you need.
Make the evaluation itself reliable
A platform records the result of your measurement design. Before using scores to approve releases:
- Define success and evidence. Check authoritative outcomes, required behavior, and prohibited side effects. Do not require one exact trajectory when several valid approaches exist.
- Validate graders. Exercise code graders with correct, incorrect, and malformed examples. Calibrate model judges against reviewed human labels; inspect false positives, false negatives, and disagreement.
- Control comparisons. Freeze dataset, scorer, model, prompt, and environment versions. Separate development cases from a held-out set, and repeat stochastic trials.
- Report uncertainty and coverage. Include sample sizes, task-category results, and appropriate uncertainty estimates. Keep infrastructure failures, missing evidence, and grader errors visible rather than treating them as passes.
- Gate critical failures explicitly. A high overall average must not erase an unauthorized payment.
For higher-is-better quality scores, p90–p99 describe the good end of the distribution. With 95 passes and five failures, all three can equal 1. Report failure counts and rates; inspect the lower tail where useful. High percentiles are appropriate for adverse quantities such as latency or cost, and no percentile replaces reviewing severe failures.
These principles complement trace inspection. The distinction between an agent's transcript and the environment's final state is explained in Anthropic's agent-evaluation guidance.
Monitoring agent performance after launch
Offline tests cover the tasks and conditions represented in the suite. Production monitoring helps identify distribution changes and missing cases. Compare three overlapping dimensions.
Scoring: code or model evaluators can assess a response, a span, a trajectory, or a conversation. Cost depends on the evidence size, number of graders, repetitions, and sampling. Reusing a rubric across offline and online evaluation helps consistency, but does not establish its validity.
Pattern discovery: Laminar's Signals analyze recorded evidence and emit structured events. Triggers, filters, and optional sampling determine coverage; the default token filter excludes traces of 1,000 tokens or fewer. Clusters organize events already detected, revealing recurring subtypes without a separate rule for each subtype. They cannot recover a failure that produced no event.
Signals are model judgments, so review flagged and unflagged traces. Event totals are not automatically failure rates: a trace may produce several events, and excluded traffic is outside the denominator. LangSmith's Insights Agent also groups production behavior, so pattern discovery is not unique to Laminar.
Intervention: runtime controls can reject or redirect detected violations before an action or response proceeds. Test their enforcement point, timeout behavior, and error handling. Offline evaluation, production analysis, and runtime enforcement can work together.
Worked example: from a failing trace to a refund regression test
Suppose a support agent times out after a payment commits, retries, and creates a second refund. The reply looks correct. The failure is in the payment outcome.
1. Collect and review the case. A Signal can flag suspicious retries for investigation. Inspect the trace and payment records. Add the relevant input to a dataset, then review it through a labeling queue. Replace the copied output with an approved reference; a failed production response is not ground truth.
2. Reproduce the environment. Each trial needs a fresh payment sandbox, the original order state, and the timeout-after-commit behavior. Route every payment tool to that sandbox. Keep the reference hidden from the agent. Export all newly committed refunds for the trial, including wrong-order payments, with one row per transaction.
3. Check the committed outcome. The following standalone grader compares the complete transaction multiset. It checks customer, order, amount in integer cents, currency, and multiplicity. Two attempts with one idempotently committed refund pass; two committed refunds fail. The trusted sandbox adapter, not the agent's reply, supplies the evidence.
from collections import Counter
FIELDS = ("customer_id", "order_id", "amount_cents", "currency")
def refund_rows(rows):
if not isinstance(rows, list):
raise ValueError("Refund records must be a list")
normalized = []
for row in rows:
if not isinstance(row, dict) or not all(k in row for k in FIELDS):
raise ValueError("Incomplete refund record")
if any(not isinstance(row[k], str) or not row[k]
for k in ("customer_id", "order_id", "currency")):
raise ValueError("Invalid refund identifier or currency")
if type(row["amount_cents"]) is not int or row["amount_cents"] <= 0:
raise ValueError("Refund amounts must be positive integer cents")
normalized.append(tuple(row[k] for k in FIELDS))
return Counter(normalized)
def refund_outcome(output, target):
if not isinstance(target, dict) or "expected_refunds" not in target:
raise ValueError("Missing reviewed reference")
expected = refund_rows(target["expected_refunds"])
if not isinstance(output, dict) or output.get("ledger_complete") is not True:
raise ValueError("Incomplete outcome evidence")
actual = refund_rows(output.get("committed_refunds"))
return int(actual == expected)
expected = {
"customer_id": "C1", "order_id": "A",
"amount_cents": 4900, "currency": "USD",
}
target = {"expected_refunds": [expected]}
assert refund_outcome(
{"ledger_complete": True, "committed_refunds": [expected]}, target
) == 1
assert refund_outcome(
{"ledger_complete": True, "committed_refunds": [expected, expected]}, target
) == 0
An approved no-refund case uses {"expected_refunds": []}. Missing references or incomplete evidence raise errors. This grader covers committed refunds only; add separate checks for authorization attempts, other side effects, and reply quality.
4. Record scores in Laminar and enforce the gate. This integration function accepts your executor and a nonempty list of reviewed cases, each containing data and target. The executor resets the sandbox, runs your agent with the case's data, and returns the evidence shape above. Call it from an ordinary Python script with the lmnr package and LMNR_PROJECT_API_KEY configured.
def evaluate_refunds(executor, cases):
from lmnr import evaluate
if not cases:
raise ValueError("The regression suite must not be empty")
scores = []
def checked_outcome(output, target):
score = refund_outcome(output, target)
scores.append(score)
return score
result = evaluate(
data=cases,
executor=executor,
evaluators={"refund_outcome": checked_outcome},
group_name="refund-agent",
)
if result is None or result.get("error_message"):
raise RuntimeError("Evaluation did not complete successfully")
if len(scores) != len(cases) or any(score != 1 for score in scores):
raise RuntimeError("Refund regression or incomplete evaluation")
return result
Run this wrapper directly as a Python script; its completion check is not designed for CLI discovery mode. Let exceptions fail CI. The count check prevents a partial run from passing, while the per-case gate avoids averaging away failures. A passing finite suite is evidence about those cases, not proof that production cannot fail.
5. Investigate changes. Keep the reference cases, ordering, and fixtures fixed across versions. Use run comparison to find changed datapoints and inspect the instrumented calls around the timeout. Include no-refund cases, failed payments, wrong amounts, extra orders, and idempotent retries in the suite.
FAQ
Which AI agent evaluation platform should I shortlist?
Laminar fits custom evaluations connected to trace debugging; Braintrust and LangSmith offer managed experiment and production-scoring workflows. Langfuse and Phoenix are candidates for running the platform yourself, with different licenses. Start with your hardest failure case and test two tools against the same evidence and grader.
How do I measure rare agent failures?
Report failure counts, severity, and rates with sample sizes and uncertainty. For quality scores where higher is better, high percentiles can conceal bad cases; they are more useful for latency or cost. Inspect individual failures alongside aggregates in your evaluation comparisons.
Can I evaluate multi-turn agents before they have users?
Yes. Use scripted or simulated user conversations to test multi-step tasks, follow-up questions, and recovery from tool failures. Define the expected outcomes and score the agent's behavior in a controlled environment. Add reviewed real interactions when available; simulated users and scenarios may differ from production.
Does production trace analysis replace offline evaluation?
No. It can expose missing cases and recurring patterns, while offline suites test changes under controlled conditions. Laminar's Signals configuration determines which traces are analyzed; validate its findings before turning them into labels or release criteria.
Is a correct tool call enough to pass an agent evaluation?
Only if tool-call correctness is the entire requirement. A refund task also needs evidence of the committed payment, correct amount and recipient, and absence of unauthorized effects. Trace inspection explains behavior; a task-specific outcome check establishes whether the intended state was reached.
Can I self-host Laminar evaluations?
Yes: tracing, datasets, labeling queues, and evaluations are available in the self-hosted platform. Signals, event clusters, and email or Slack alerts require an enterprise license for self-hosted deployments. See the deployment comparison for feature boundaries.