Skip to main content
Laminar has a plugin for Harbor, so the runs you already launch with harbor run show up in Laminar as evaluations. Every trial gets its own trace with the agent’s LLM calls and tool calls, the verifier’s rewards become scores, and runs on the same dataset are charted together so you can see whether a new agent or model moved the score. Harbor is a framework for evaluating agents in sandboxed environments: it runs an agent such as Claude Code, Codex, or your own on a dataset like Terminal-Bench, then runs each task’s verifier to score the result. The Laminar plugin hooks into Harbor’s job and trial events. You don’t change your agent or your tasks.

Getting started

1

Install the Laminar SDK next to Harbor

Install lmnr in the same Python environment as harbor:
If you installed Harbor with uv tool install harbor, add lmnr to that tool environment instead:
Check that Harbor picked up the plugin:
The output lists laminar with the import path lmnr.integrations.harbor:LaminarPlugin.
2

Set your project API key

Get the key from Project Settings → Project API keys in Laminar.
3

Run Harbor with the plugin

Add --plugin laminar to your usual command:
The plugin logs a link to the evaluation when the job starts. Trials appear in Laminar as they finish.

Agents written in TypeScript or any other language

Harbor loads plugins from Python, so the Laminar plugin ships in the Python SDK (lmnr), not in @lmnr-ai/lmnr. That doesn’t limit which agents you can evaluate. The plugin reads Harbor’s trial results and the agent’s ATIF trajectory, not the agent process, so it reports any agent Harbor can run:
  • Node.js and TypeScript agents such as Claude Code, Gemini CLI, OpenCode, and Harbor’s langgraph agent on its Node runtime.
  • Python agents, and agents in any other language Harbor can run in its sandboxes.
For a TypeScript agent, you still install lmnr with pip next to harbor. Nothing gets installed inside the sandbox.

What you see in Laminar

Each Harbor job becomes one evaluation, and each trial becomes one datapoint in it. Every trial is its own trace. The root EVALUATION span holds the trial’s rewards and, if the trial failed, the exception. Under it are four spans for the trial’s phases, with Harbor’s own timings:
  • environment_setup
  • agent_setup
  • agent, an EXECUTOR span with the agent’s total tokens and cost.
  • verifier, an EVALUATOR span whose output is the rewards.
When the agent writes an ATIF trajectory (agent/trajectory.json in the trial directory), the plugin replays it under the agent span. Every model turn becomes an LLM span with the model, provider, messages, tokens, and cost. Every tool call becomes a TOOL span with its arguments and result. Subagent trajectories are nested under the tool call that started them. Open a datapoint and you read the agent’s run as a transcript, not a list of span names. A few details on how results are counted:
  • Retries overwrite the same datapoint, so a trial Harbor retried shows the final attempt.
  • Trials without rewards (for example, the agent crashed before the verifier ran) score 0. Harbor’s mean counts them the same way, so the averages in Laminar match Harbor’s.
  • Cancelled trials get no scores and are marked cancelled in metadata.
  • Errors talking to Laminar are logged and don’t stop the Harbor run, unless you set fail_fast.

Options

Pass options with --pk key=value on harbor run. Most options can also be set with an environment variable. For example, to compare two agents in one group:
If you pass more than one --plugin, target the Laminar plugin with --pk laminar.group_name=terminal-bench-agents.

Self-hosted Laminar

Point the plugin at your instance with base_url, http_port, and grpc_port, the same settings as self-hosted evaluations:

Next steps

Read a failing trial

Open a low-scoring datapoint and read the agent’s run in the transcript view to see where it went wrong.

Compare runs

Runs in the same group share a progression chart, and you can diff two runs side by side.

Track failure patterns

Define a Signal such as “the agent gave up before running the tests” and count it across every trial.

Query across trials

Ask which tasks fail for every agent from the SQL editor, the SQL API, or the MCP server.

Claude Code

Trace Claude Code sessions with a one-command plugin.

Codex

Trace Codex CLI sessions with a one-command plugin.

OpenCode

Trace OpenCode coding agent sessions.

Claude Agent SDK

Trace Claude Agent SDK runs and their subagents.

LangChain / LangGraph

Trace LangChain chains and LangGraph graphs.

All integrations

Browse every provider, framework, coding agent, and browser integration Laminar supports.