harbor run show up in Laminar as evaluations. Every trial gets its own trace with the agent’s LLM calls and tool calls, the verifier’s rewards become scores, and runs on the same dataset are charted together so you can see whether a new agent or model moved the score.
Harbor is a framework for evaluating agents in sandboxed environments: it runs an agent such as Claude Code, Codex, or your own on a dataset like Terminal-Bench, then runs each task’s verifier to score the result. The Laminar plugin hooks into Harbor’s job and trial events. You don’t change your agent or your tasks.
Getting started
1
Install the Laminar SDK next to Harbor
Install If you installed Harbor with Check that Harbor picked up the plugin:The output lists
lmnr in the same Python environment as harbor:uv tool install harbor, add lmnr to that tool environment instead:laminar with the import path lmnr.integrations.harbor:LaminarPlugin.2
Set your project API key
3
Run Harbor with the plugin
Add The plugin logs a link to the evaluation when the job starts. Trials appear in Laminar as they finish.
--plugin laminar to your usual command:Agents written in TypeScript or any other language
Harbor loads plugins from Python, so the Laminar plugin ships in the Python SDK (lmnr), not in @lmnr-ai/lmnr. That doesn’t limit which agents you can evaluate. The plugin reads Harbor’s trial results and the agent’s ATIF trajectory, not the agent process, so it reports any agent Harbor can run:
- Node.js and TypeScript agents such as Claude Code, Gemini CLI, OpenCode, and Harbor’s
langgraphagent on its Node runtime. - Python agents, and agents in any other language Harbor can run in its sandboxes.
lmnr with pip next to harbor. Nothing gets installed inside the sandbox.
What you see in Laminar
Each Harbor job becomes one evaluation, and each trial becomes one datapoint in it.
Every trial is its own trace. The root
EVALUATION span holds the trial’s rewards and, if the trial failed, the exception. Under it are four spans for the trial’s phases, with Harbor’s own timings:
environment_setupagent_setupagent, anEXECUTORspan with the agent’s total tokens and cost.verifier, anEVALUATORspan whose output is the rewards.
agent/trajectory.json in the trial directory), the plugin replays it under the agent span. Every model turn becomes an LLM span with the model, provider, messages, tokens, and cost. Every tool call becomes a TOOL span with its arguments and result. Subagent trajectories are nested under the tool call that started them. Open a datapoint and you read the agent’s run as a transcript, not a list of span names.
A few details on how results are counted:
- Retries overwrite the same datapoint, so a trial Harbor retried shows the final attempt.
- Trials without rewards (for example, the agent crashed before the verifier ran) score
0. Harbor’s mean counts them the same way, so the averages in Laminar match Harbor’s. - Cancelled trials get no scores and are marked
cancelledin metadata. - Errors talking to Laminar are logged and don’t stop the Harbor run, unless you set
fail_fast.
Options
Pass options with--pk key=value on harbor run. Most options can also be set with an environment variable.
For example, to compare two agents in one group:
--plugin, target the Laminar plugin with --pk laminar.group_name=terminal-bench-agents.
Self-hosted Laminar
Point the plugin at your instance withbase_url, http_port, and grpc_port, the same settings as self-hosted evaluations:
Next steps
Read a failing trial
Open a low-scoring datapoint and read the agent’s run in the transcript view to see where it went wrong.
Compare runs
Runs in the same group share a progression chart, and you can diff two runs side by side.
Track failure patterns
Define a Signal such as “the agent gave up before running the tests” and count it across every trial.
Query across trials
Ask which tasks fail for every agent from the SQL editor, the SQL API, or the MCP server.
Related integrations
Claude Code
Trace Claude Code sessions with a one-command plugin.
Codex
Trace Codex CLI sessions with a one-command plugin.
OpenCode
Trace OpenCode coding agent sessions.
Claude Agent SDK
Trace Claude Agent SDK runs and their subagents.
LangChain / LangGraph
Trace LangChain chains and LangGraph graphs.
All integrations
Browse every provider, framework, coding agent, and browser integration Laminar supports.