The first thing you usually want to know about an agent run is what it was asked to do. The answer is in the user's message, but rarely on its own. By the time the message reaches the model, the agent's harness has wrapped it in timestamps, environment state, system notices and user details, and every agent wraps it differently. This post is about how we find the task inside all that: cheaply, across millions of runs a day, and without asking developers where to look.
Why Tasks?
Laminar processes millions of agent runs every day, and every one of them starts with a task. By task we mean the request as the user wrote it, verbatim, not a summary or a reconstruction of intent. It's one of the most useful things to know about a run. With the task in hand, you can:
- see at a glance what an agent was asked to do when you open its trace
- understand what your users actually ask for
- find which tasks an agent fails at
- build evals for different kinds of tasks from production data
The Challenge
The obvious answer is to take the role: "user" messages from the agent's first LLM call. Sometimes that's enough: many agents send the task as a plain user message of its own.
But just as often, the task is buried among other context. One common reason is prompt caching: cache hits need a stable prefix, so many agents keep the system prompt fixed and move everything that changes per run into the user messages. The user turn ends up being a mix of:
- static scaffolding: instructions, section headers and markers
- dynamic request context: who sent the message, the pull request under review, a customer's profile
- dynamic environment state: the current time, open browser tabs, files on disk, the agent's todo list
- the task itself: usually the one part that's different in every run
For example, here is the first LLM call of a coding agent's run. After the system prompt, the user turn arrives as three separate user messages:
Only one line of this is the task: "The checkout button still doesn't work on mobile." The rest is scaffolding: the state of the repository, a system notice, who sent the message, and an instruction for the reply. And the scaffolding isn't even fixed, since the repository state changes on every turn.
We call these three messages together the user turn: the role: "user" messages that come after the agent's last reply. In a new conversation, that's all of them. In a follow-up, the LLM call also carries the earlier conversation, and the user turn is just the new messages since the agent last replied. The user turn is where we look for the task.
For the agent's developer, none of this is hard: they know exactly where the task is, and a custom UI renderer or a quick SQL query can pull it out. But that stops working once the agent changes, and developers change prompts every day. They add and remove sections, rename XML tags, and move context from one message to another. Every change can silently break a hand-written extraction. And we need to extract tasks for every agent on the platform, with no developer to ask.
TL;DR: an LLM writes a task extraction regex for each template, the format an agent wraps around every task, so most runs need no LLM call at all. When the format changes, we notice by comparing each run with similar recent runs, and write a new regex for the new format. The rest of this post walks through the attempts that led to this design.
Attempt #1 – Just Ask an LLM
The simplest solution is to ask an LLM to extract the task from every run. Understanding the task of a single run is easy enough that even small models handle it well.
The problem is cost. At a million agent runs per day and about 2,000 tokens per user turn, that's roughly $6,000 a month, and it grows linearly with traffic. We want the task in every trace of every agent by default, and at that price we can't afford it.
Attempt #2 – Regexes
Think about how you'd find the task in a trace of your own agent. You wouldn't read the whole thing; you'd jump straight to where you know it is: right after the system reminder, after the USER TASK: line, or the last message of the turn.
That's a pattern, and a pattern doesn't need an LLM on every run. It can be a regex. So instead of asking an LLM for the task, we ask it once per agent for a regex that finds the task, and then apply that regex to every following run without another LLM call.
Writing that regex is harder than reading one turn, so we give the model several samples from the same agent side by side. What stays the same between them is scaffolding; the task is somewhere in what changes, and the model decides which part it is. Sometimes the right answer is simply "take everything" ((?s)(.*)), since plenty of agents send the task as a bare message.
For the coding agent above, the regex takes whatever sits between </system_reminder> and <user_info>:
This works well until the agent changes. Nobody writes a template once and never touches it again, and a changed template can break a regex in two ways:
-
Silently. The developer starts attaching files to the user's message:
... The checkout button still doesn't work on mobile. <attachments> - mobile-checkout.png (screenshot, 1.2 MB) </attachments> <user_info> ...The regex still matches, so nothing fails, but every extracted "task" now ends with a list of attachments.
-
Loudly. The developer moves <user_info> into the first message, next to the repository state. Now nothing follows the task, the regex finds no <user_info> after the system reminder, and every run comes back with no task at all.
Either way, the problem changes from "what's the task of this run?" to "how do we know that the pattern we have is no longer valid?" In other words: which version of the template does this run use, and when did it change?
Attempt #3 – Parsing the Structure
As the examples show, the task is usually wrapped in some kind of marker. Prompts are written to be clear to a model, so they tend to have a visible structure – often the whole user turn is built from blocks with dynamic content inside. And very often, those blocks are XML tags.
So the first idea is simple: treat the XML tags as the prompt's structure. Find all of them in the user turn in the order they appear, join and hash them, and treat the result as the template version. When the developer adds, removes or renames a tag, the version changes, and we generate a new regex from the latest samples.
That catches the attachments change from the previous section: the new <attachments> tag changes the list of tags, so the version changes, and the new version gets a new regex:
It's cheap, and for many agents it's enough. But it breaks in both directions:
-
Too sensitive. XML tags are not always section markers. A browser agent may include the page's interactive elements (<button>, <input>, <a>) in its user turn, and a coding agent may include HTML from the codebase. The signature then changes on almost every run, the regex is regenerated every time, and the cost goes back up.
-
Too blind. Many templates don't use XML at all: an agent may get its context as plain text, under Markdown headers, or with any other marker the developer likes. Say the developer had added the attachments as plain text instead:
The checkout button still doesn't work on mobile. Attachments: - mobile-checkout.png (screenshot, 1.2 MB) <user_info> ...The tag list is exactly what it was, so the version doesn't change, and the old regex keeps returning the attachments as part of the task.
The core problem is that we're trying to understand the structure of a prompt, and there is no universal structure to understand.
Attempt #4 – Learning from Many Runs
So instead of parsing the structure, we learn the template from the runs themselves.
Looking at a single run, you can't tell what is static and what is dynamic; there's simply not enough information. But looking at many runs of the same version, it becomes obvious: the static part is whatever they all share.
We treat each user turn as a sequence of lines, and define a version as the lines that all its turns have in common, in the same order: their longest common subsequence (LCS). Here it is for three runs of our coding agent, with the attachments added as plain text, as above:
The version is a hash of those common lines. No assumptions about XML, Markdown or any other markers: a dynamic line is simply one that differs between runs. The task is almost always dynamic, since it differs from run to run. The LCS only tells us which version a run belongs to; finding the task in the run is still the regex's job.
This also catches the change that fooled Attempt #3. Attachments: is the same in every run that has it, so it is a static line, and adding it changes the version. The file name under it differs from run to run, so it's dynamic and never affects the version – which is right, because it's content, not template.
In practice, we compute this over a window of recent runs: version = LCS(recent runs). If the result matches a known version, we reuse its regex; if not, it's a new version and gets a new regex. When the template changes, the first runs of the new version still share the window with the old one, so their version is only what the two have in common. As more new runs arrive, the window resolves to the new version on its own, after a short lag. During the lag, runs keep the old version's regex.
This has a hidden assumption, though: that recent runs all come from the same version. But what if a developer runs an A/B test?
Final Design – Top-K Nearest Neighbors
During an A/B test, runs arrive as A B A B A B …. Every window contains both versions, so the LCS is only what the two share, forever. At least one of the versions the developer actually shipped never shows up.
And A/B tests aren't the only case. The coding agent from our first example gets chat turns from its users, but also long automated requests from a CI integration, like "review this pull request". That's two different templates, interleaved all day.
The fix: for each new run,
- Compare it with the runs of the last hour, by the share of lines their user turns have in common with it.
- Select the closest X% of them, say 30%. If there are other A runs in the window, an A run's closest neighbors will be A runs.
- Run the LCS only on those.
So the algorithm becomes version = LCS(X% closest of the last hour's runs).
Choosing the numbers. The window is a time span, not a count: a fixed "last 100 runs" could all come from one user session on a busy agent, and that user's name would look static; an hour of traffic almost never does. It's also capped at a couple hundred distinct turns per agent, so ranking it stays cheap. K is a share rather than a count, so it scales with the window: with 30%, a run's closest neighbors come from its own variant even in a 50/50 A/B test.
Versioning itself needs no LLM at all, only hashing and LCS. The LLM comes in once per version: each extraction regex belongs to an agent and a template version, and when a new version appears, we generate its regex from fresh samples of that version. Until it's ready, we fall back to an LLM per run. The result: LLM cost grows mostly with how often you change your agent, not with how much traffic it gets. Today we spend around $10 a day on task extraction, instead of the $200 a day it would take to call an LLM on every run.
A Few More Details
- Which agent? A trace often holds several agents: the main one, its subagents, and small helpers that summarize the conversation or write a title. We tell agents apart by the first sentence of their system prompt, and extract the task for the main one, whose first LLM call has the most input tokens.
- First turns and follow-ups. Follow-up turns usually follow a different template than the first turn of a conversation, so they're versioned separately.
- Message boundaries. We join a turn's messages with a marker line, so a regex can anchor on where one message ends. The marker is stripped from whatever the regex captures.
What's Still Hard
There are still edge cases. Very long turns are truncated before we generate a regex, so extraction can miss on them. Large context that barely changes between runs, or tasks that repeat verbatim like "Continue", can look static and split one template into extra versions, which costs a few more LLM calls. A variant smaller than the share, like a 10% rollout, still gets mixed in with the others. And a quiet agent needs a couple dozen runs before it gets a version.
We tried to keep such cases rare while keeping costs low. An explicit signal from the developer, like a CI step that tells Laminar when the template changed, would close the gap, but we didn't want to ask users for extra work.
Wrapping Up
You can't learn a template by looking at one turn, but you can by looking at many: treat turns as lines, compare each one to its closest neighbors, and accept a short lag in exchange for a pipeline that corrects itself.
This is what powers task extraction in Laminar today. The agent's task is in every trace, and it's also a column you can query: agent_input on the traces table. One query in the SQL editor shows what your agents were asked to do, or which tasks end in errors. There's nothing to configure, and it keeps up as your agent changes.
Try Laminar and see what your agents are being asked to do.