Skip to content

Article

Evaluating AI Agents: Building Tasks That Grade Them Fairly

Build one coding-agent task from instruction to verifier, then test it for loopholes, flakes, reward hacking, and repeatable results.

Published 29 Sept 2026Updated 1 Oct 202616 min read
Agent Evaluation · Harbor · Benchmarks · Reward Hacking · Rubrics · LLM-as-Judge
An agent task flowing into an outcome verifier, followed by oracle, no-op, and cheating-agent checks with their expected scores.
On this page (21)

In August 2024, OpenAI had 93 software developers review 1,699 tasks from SWE-bench, one of the most widely used benchmarks for coding agents. About 539 survived the filtering process, and 500 were sampled for SWE-bench Verified. Some tasks described the problem too vaguely to solve; many more had tests that "may unfairly mark valid solutions as incorrect." In February 2026, OpenAI audited the Verified tasks that its o3 model kept failing and found that at least 59.4% had flawed tests that "reject functionally correct submissions."

So when someone says a new agent scores 62%, there are two questions behind the number. Do the tasks measure what they claim to measure? And would the agent score 62% again tomorrow? Most of the work in evaluating agents is making the answer to both questions yes, and it happens long before anyone looks at a leaderboard.

At Mercor I work on this part-time. I write agent tasks in Harbor, an open-source framework for running and grading agents; run agents on them and read what they did; review tasks other authors wrote; and write rubrics for work that can't be checked by a test. None of that work appears here. The task this chapter follows is one I made up for it, and every number comes from a public source.

Five rules shape the rest of this chapter:

  1. Where the outcome can be checked by running something, grade it with a test. A test is cheaper and faster than a judge, and much harder to argue with.
  2. Don't trust a task until three checks pass: the reference solution scores full marks, doing nothing scores zero, and an agent told to cheat scores zero.
  3. Run every task several times, and report how often the agent succeeds every time, not only whether it ever did.
  4. Read the runs. A score says that something happened; the run says what.
  5. Use rubrics only for what can't be tested, and write each criterion so that two experts would give the same answer.

How to read the diagrams

  • The task: instruction, environment, files
  • The agent and what it did
  • The verifier and passing results
  • Reference solutions and task checks
  • Failures: cheats, flaky or unfair tests

The words this chapter uses

Evaluation has its own vocabulary, and most of the confusion in discussions about agents comes from mixing these terms up. The running example comes in the next section; this table shows each term as it applies there.

terms
term
agent
meaning
A model in a loop that reads, runs commands, edits files, and decides its next step
in the running example
A coding agent working inside a container
term
task
meaning
One problem to solve: an instruction, an environment to solve it in, and a way to check the result
in the running example
Fix a log shipper that loses lines when its log file rotates
term
trial (run)
meaning
One agent's attempt at one task
in the running example
One attempt, start to finish, in a fresh container
term
trajectory
meaning
The record of a trial: every command, file read, edit, and message
in the running example
The agent read shipper.py, rotated the log by hand, then edited the loop
term
verifier
meaning
The program that checks the result after the agent stops
in the running example
A test that rotates the log mid-stream and compares the shipped lines
term
reward
meaning
The number the verifier writes: usually 1 for pass, 0 for fail
in the running example
1 if all 500 lines arrived exactly once and in order
term
reference solution
meaning
A known-good answer, used to prove the task can be solved. Harbor runs it with an agent called oracle
in the running example
A shell script that installs a correct shipper
term
rubric
meaning
A list of criteria for grading output that no test can check, such as an explanation
in the running example
Criteria for the pull request description

Two kinds of task

Some outcomes can be checked by running something: the service starts, the numbers match, the file contains exactly the right lines. Others can't: whether an explanation is correct and complete, whether a code review caught the real problem. The two need different graders, and they fail in different ways.

Graded by a test

Verifiable tasks

Grader
A script that inspects the final state
Cost per grade
Seconds of compute
Main risk
Tests that reject correct answers or accept cheats
Example
Every log line arrives exactly once after rotation

Graded against criteria

Rubric-graded tasks

Grader
An expert, or a model checked against experts
Cost per grade
A model call per criterion, or an expert's time
Main risk
Criteria that two graders read differently
Example
The pull request explains the real cause

Terminal-Bench, a benchmark of hard tasks that agents solve in a Linux terminal, is built entirely from the first kind. Its paper says the tests "do not test the agent's commands or console output. This is intentional as Terminal-Bench is an outcome-driven framework." Anthropic's guide to agent evals makes the same point from the other direction: "it's often better to grade what the agent produced, not the path it took." Where both kinds apply, as they do in the running example, the test decides whether the work is correct, and the rubric grades what the test can't see.

The running example: a log shipper that loses lines

A small service, shipper.py, follows an application's log file and copies every new line to an output file, standing in for a real log pipeline. Every night the log is rotated the way logrotate does it by default: the current file is renamed to app.log.1, and a new, empty app.log is created.

The shipper opens the log once and reads it forever. It never notices the rename:

How the shipper loses lines at rotation

  1. 02:59:58

    Reading app.log normally

    The shipper holds an open file handle. On Linux, the handle points at the file itself, which the filesystem identifies by an inode number, not at the name.

    app.log → inode 4711 · shipped: lines 1–100

  2. 03:00:00

    logrotate renames and recreates

    app.log becomes app.log.1, still inode 4711. A new, empty app.log is created with inode 4712.

    app.log.1 → 4711 · app.log → 4712

  3. 03:00:01

    The application keeps writing

    The application opens the log by name, so lines 101 onward go into the new file.

    inode 4712: lines 101, 102, 103 …

  4. 03:00:01 →

    The shipper waits on a file nobody writes to

    Its handle still points at inode 4711. It reaches the end of that file and waits for more lines forever. Nothing after line 100 is ever shipped.

    shipped: lines 1–100 · lost: everything after

A correct fix has two parts. At the end of the file, the shipper checks whether the name app.log now points at a different inode; if it does, the old file has been rotated away, so it opens the new one and reads it from the start. And before it lets go of the old file, it reads it to the end one more time, because the application may have written a few lines to it between the shipper's last read and the rename. It's easy to get the first part right and miss the second, and that gap is what the task is really testing.

Anatomy of a Harbor task

Harbor is an open-source framework "from the creators of Terminal-Bench for evaluating and optimizing agents and language models," and "the official harness for Terminal-Bench-2.0." Its documentation puts a task in one sentence: "A Harbor task is an instruction, environment, and test script." On disk, the running example looks like this:

Sketch: the fix-log-shipper task
fix-log-shipper/
├── instruction.md        what the agent is asked to do
├── task.toml             timeouts, resources, free-form metadata
├── environment/
│   ├── Dockerfile        the container the agent works in
│   └── shipper.py        the buggy service
├── solution/
│   └── solve.sh          the reference solution (run by the oracle agent)
└── tests/
    ├── test.sh           the verifier: writes the reward
    └── test_rotation.py

Each file has a job, and most task bugs come from one of them doing another's job.

The instruction

Sketch: instruction.md
The log shipper at /app/shipper.py follows /var/log/app/app.log and appends
every line to /logs/artifacts/shipped.log. When the log is rotated, it loses lines.
 
Fix /app/shipper.py so that every line written to /var/log/app/app.log reaches
/logs/artifacts/shipped.log exactly once and in order, across any number of rotations.
Rotation works like logrotate's default: the current file is renamed to
app.log.1 and a new, empty app.log is created. Keep the output format: one
log line per output line, unchanged. The verifier performs rotations at least
one second apart and may write unread lines immediately before each rename.

Every behaviour the tests will check is stated here: exactly once, in order, any number of rotations, the rotation method, the output format. Harbor's guidance is to "use absolute environment paths (/app/out.json not out.json) and define schemas for the output so instructions and tests agree." One detail surprises new authors: "the instruction is not passed to the agent as a file in the environment." The agent receives it as text, so it can't be found by searching the container.

The configuration

Sketch: task.toml
[metadata]                 # free-form: Harbor doesn't enforce these fields
author_name = "Robin Singh"
difficulty = "medium"
category = "debugging"
 
[agent]
timeout_sec = 900          # without this, Harbor enforces no agent timeout
 
[verifier]
timeout_sec = 600
 
[environment]
build_timeout_sec = 600
cpus = 1
memory_mb = 2048
network_mode = "no-network"

The metadata table is, in the documentation's words, "arbitrary metadata provided by the task author"; difficulty and category are conventions, not schema. The timeouts and resources matter more than they look. Terminal-Bench's own maintainers wrote that "timeouts affect task difficulty," and later fixed eight tasks whose resource budgets were too small for at least one valid approach to finish reliably.

The environment

Sketch: environment/Dockerfile
FROM python:3.12-slim
WORKDIR /app
COPY shipper.py /app/shipper.py
RUN mkdir -p /var/log/app /logs/artifacts

What's missing matters as much as what's there. The tests and the reference solution are not copied into the image. Harbor's own task checker asks exactly this: "Are the tests/ folder or solution/ folder copied to the image? They should not be."

The verifier

Sketch: tests/test.sh
#!/bin/bash
# Runs after the agent has stopped, in the same container.
mkdir -p /logs/verifier
pip install --quiet pytest==8.4.1
if pytest -q /tests/test_rotation.py; then
  echo 1 > /logs/verifier/reward.txt
else
  echo 0 > /logs/verifier/reward.txt
fi
Sketch: tests/test_rotation.py
import os, signal, subprocess, time
from pathlib import Path
 
LOG = Path("/var/log/app/app.log")
OUT = Path("/logs/artifacts/shipped.log")
 
def rotate():
    os.rename(LOG, LOG.with_name("app.log.1"))  # logrotate's default: rename…
    LOG.touch()                                  # …then create a new, empty file
 
def wait_for_lines(n, deadline=20.0):
    end = time.monotonic() + deadline
    while time.monotonic() < end:
        if OUT.exists() and len(OUT.read_text().splitlines()) >= n:
            return
        time.sleep(0.05)
    raise AssertionError(f"timed out waiting for {n} shipped lines")
 
def append_lines(first, last):
    with LOG.open("a") as f:
        for i in range(first, last + 1):
            f.write(f"line {i}\n")
        f.flush()
        os.fsync(f.fileno())
 
def test_every_line_arrives_exactly_once_in_order():
    OUT.unlink(missing_ok=True)
    LOG.write_text("")
    shipper = subprocess.Popen(["python3", "/app/shipper.py"])
    try:
        start = 1
        for boundary in (100, 250, 400):
            append_lines(start, boundary - 3)
            wait_for_lines(boundary - 3)
 
            # Freeze the reader, append an unread tail to the old file,
            # rotate, then resume. A correct shipper drains the old handle.
            os.kill(shipper.pid, signal.SIGSTOP)
            append_lines(boundary - 2, boundary)
            rotate()
            os.kill(shipper.pid, signal.SIGCONT)
            wait_for_lines(boundary)
            time.sleep(1.0)
            start = boundary + 1
 
        append_lines(start, 500)
        wait_for_lines(500)
        time.sleep(0.5)  # give late duplicates a chance to show up
        assert OUT.read_text().splitlines() == [f"line {i}" for i in range(1, 501)]
    finally:
        shipper.kill()

The test starts the shipper, writes 500 lines, and controls the race around three rotations. SIGSTOP freezes the process; three unread lines are appended to the old file, the file is renamed, and SIGCONT resumes the reader. The wait_for_lines handshake proves each rotation completed before the next one. Comparing the complete list catches loss, duplication, reordering, and altered text. The script then writes Harbor's numerical reward under /logs/verifier/.

The reference solution

Sketch: solution/solve.sh
#!/bin/bash
cat > /app/shipper.py <<'EOF'
import os, time
from pathlib import Path
 
LOG, OUT = "/var/log/app/app.log", "/logs/artifacts/shipped.log"
 
def open_log():
    while True:
        try:
            f = open(LOG)
            return f, os.fstat(f.fileno())
        except FileNotFoundError:
            time.sleep(0.05)
 
Path(OUT).parent.mkdir(parents=True, exist_ok=True)
f, opened_stat = open_log()
with open(OUT, "a") as out:
    while True:
        line = f.readline()
        if line:
            out.write(line)
            out.flush()
            continue
        try:
            rotated = not os.path.samestat(os.stat(LOG), opened_stat)
        except FileNotFoundError:
            rotated = False
        if rotated:
            for line in f:          # lines written just before the rename
                out.write(line)
            out.flush()
            f.close()
            f, opened_stat = open_log()  # the new file, from its first line
        else:
            time.sleep(0.05)
EOF

The solution folder is, per the documentation, "a reference script used by the oracle agent to sanity-check that a task is solvable." It's optional for published benchmarks, but "without solution/, the Oracle agent cannot run," and without the oracle you have no proof that your own task can be passed.

What happens when you run it

A run in Harbor is called a trial: "one agent's attempt at completing one task." A job is a collection of trials, run in parallel. The order of events inside one trial explains most of the design above:

One trial, from command to reward

  1. harbor run

    The command names the task, the agent, the model, and how many attempts to make.

    harbor run -p ./fix-log-shipper -a claude-code -m <model> -k 5

  2. The environment is built

    The container is built from environment/, within the build timeout.

    Dockerfile → image · build_timeout_sec 600

  3. The agent works

    The agent receives the instruction as text and works in the container until it stops or hits its timeout. The tests and the solution are not in the container.

    agent/trajectory.json records every step

  4. The tests are uploaded

    Only now does Harbor copy tests/ into the container.

    tests/ → /tests

  5. The verifier runs

    test.sh runs within its own timeout and writes the reward.

    /logs/verifier/reward.txt → 1 or 0

  6. The reward is collected

    Harbor reads the reward file. A missing or non-numeric reward is an error, not a zero, and by default isn't retried.

    result.json · /logs/artifacts downloaded · container deleted

  7. The job is summarised

    Rewards are averaged across tasks; missing rewards count as 0. With binary rewards and several attempts, Harbor also reports pass@k.

    harbor view ./jobs to read the trials

The detail that matters most is the fourth step. In the documentation's words, "the tests/ directory is uploaded to the environment at /tests/ after the agent runs." An agent can't read tests that don't exist yet. Harbor can go further and run the verifier in a separate container, where "the agent's filesystem changes are not inherited." Files deliberately published under /logs/artifacts/, plus any paths declared as artifacts, are copied into that verifier. The running example puts shipped.log there so the isolated verifier can inspect the outcome without inheriting the rest of the agent's filesystem.

Three checks before anyone trusts a task

A task is software, and like software it has bugs. Terminal-Bench's continuous-integration checks for every task are a good minimum, and all three can run before any real agent touches the task:

task checks · fix-log-shipperexpected results; the cheat is illustrative
check
oracle
what it proves
The task can be solved, and the tests accept a correct answer
command
harbor run -p ./fix-log-shipper -a oracle
must score
1
check
nop (does nothing)
what it proves
The tests don't pass the starting state: the bug is real and the tests catch it
command
harbor run -p ./fix-log-shipper -a nop
must score
0
check
cheating agent
what it proves
An agent told to game the grader can't: no reading answers, no editing the checker
command
an agent prompted to reward-hack
must score
0

The Terminal-Bench paper describes the same gate for its 2.0 release: "an automated workflow ran the task's oracle solution to ensure solvability," and "a no-op 'dummy' agent should fail the task." Each of its 89 tasks was then "manually verified by three human reviewers for correctness," using about three hours of combined reviewer attention on average. Harbor packages a first pass of that review as a command: harbor check <task-dir> grades a task against a rubric of common problems, including whether the tested behaviour is described in the instruction, whether Python dependencies are pinned, and whether the solution is hard-coded.

Tests that grade the result, not the method

A test can fail a task in two directions. It can reject a correct answer, or it can accept a wrong one. OpenAI's 2026 audit of SWE-bench Verified sorted the flawed tests it found into "narrow" tests, which reject valid solutions, and "wide" tests, which check for behaviour the task never asked for.

SWE-bench Verified · what the reviews found
review
August 2024: 1,699 original tasks, each labelled by three developers
finding
38.3% flagged for underspecified problem statements; 61.1% for unit tests that may unfairly mark valid solutions as incorrect; 68.3% filtered out, leaving 500
review
February 2026: 138 Verified tasks that OpenAI's o3 didn't solve consistently over 64 runs, each reviewed by at least six engineers
finding
59.4% had flawed tests that reject correct submissions: 35.5% narrow, 18.8% wide, 5.1% other
review
The narrow-test example
finding
A test that imports a function called get_annotation, a name the problem description never mentions. Any fix that names the function differently fails

The same thing is easy to do in the running example. Suppose a task author, trying to be thorough, adds a test that the fix reads st_ino directly. That is one way to recognise a rename, but it is not the only correct one. Python's os.path.samestat compares the device and inode from two stat results; an inotify watcher could also report the move. A test tied to one implementation rejects the others:

two correct fixes, graded by two tests
fixnarrow test: “uses st_ino”outcome test: “500 lines, once, in order”
Fix A: compares inode numbers
checks os.stat(LOG).st_inopasspass
Fix B: compares complete file identities
uses os.path.samestat on the open handle and current pathfail: wronglypass

Terminal-Bench's criterion for this is the cleanest statement I know: "the unit tests will pass if and only if the container ends in an acceptable state." Its guidance for task authors says the same in practical terms: "Tests should validate outcomes rather than implementation details. They should allow alternate correct solutions, avoid brittle source-code checks, and protect against reward hacking or accidental leakage of the reference answer."

Wide tests are the mirror image, and they usually come from the instruction and the tests drifting apart. When the Terminal-Bench maintainers fixed 28 of their 89 tasks for version 2.1, one example was a task whose "tests expected Spark SQL output, while the instructions asked for PostgreSQL." The fix is procedural: every behaviour the tests check must be written in the instruction, and every requirement in the instruction must be checked by a test.

When the agent games the test

Agents trained with reinforcement learning are rewarded for passing checks, and some learn to pass the check instead of doing the task. This is called reward hacking, and it's well documented:

Published cases of agents gaming their graders

  1. Feb 2025

    Anthropic: special-casing tests

    The Claude 3.7 Sonnet system card: the model “occasionally resorts to special-casing in order to pass test cases,” most often “directly returning expected test values rather than implementing general solutions, but also includes modifying the problematic tests themselves.” Anthropic's later Sonnet 4.5 system card reported much lower rates on updated reward-hacking evaluations, while noting that hard-coding and special-casing still occurred.

    Claude 3.7 and Sonnet 4.5 system cards

  2. Mar 2025

    OpenAI: “hack verify to always return true”

    Reasoning traces from coding tasks included plans like “Test only calls analyze.prove and verify. They don't inspect details. So we can hack verify to always return true.” The accompanying paper lists exiting early with exit(0), raising SkipTest, and writing stubs where test coverage is thin.

    OpenAI, Mar 2025; arXiv 2503.11926

  3. Jun 2025

    METR: patching the scorer

    o3 “decides to patch the competition evaluation function so that it judges every submission as successful.” On one task it hacked in every run. Telling the model not to cheat “had a nearly negligible effect.”

    METR, Recent Frontier Models Are Reward Hacking

  4. Jan 2026

    Terminal-Bench: an agent built to find exploits

    The benchmark's authors ran an adversarial agent against their own tasks. It found “monkey-patching test environments, guessing answers, and generating all possible answers with the hope that tests might only check existence of a correct answer but not absence of wrong answers.”

    arXiv 2601.11868, appendix B.4

  5. Apr 2026

    Terminal-Bench leaderboard submissions

    The integrity review found different failures: OB-1 had modified timeouts in its original Terminal-Bench 1.0 submission; Pilot uploaded the tests folder; and ForgeCode runs downloaded solutions into AGENTS.md. The maintainers said reward-hacked trials would score 0 and announced an agent judge for passing trials.

    tbench.ai, Leaderboard Integrity Update

Anthropic notes that its own measurements come from evals "explicitly designed to stress-test hacking propensities" and don't reflect real-world rates, so these cases say that hacking happens, not how often. For a task author, that's enough: any gap in the grading will eventually be found.

The Terminal-Bench exploit list maps directly onto the running example. A shipper that simply writes every line twice contains all 500 expected lines. A test that only checks that each expected line appears somewhere would pass it:

a cheating shipper against two tests
outputweak test: “every line appears”strict test: “exactly this list”
Correct shipper
500 lines, once each, in orderpasspass
Shipper that ships every line twice
1,000 lines: each expected line, twicepass: wronglyfail

The defences are mostly about where things live and what the test trusts. Keep answers out of reach: Harbor uploads the tests only after the agent finishes, the solution never goes into the image, and a task can cut the agent off from the network during its run, which blocks the most common loophole on the Terminal-Bench leaderboard. Check for wrong answers as well as right ones, by comparing the exact output, so duplicates and extra lines fail. This test writes its own log lines and starts the shipper itself instead of asking the agent's code whether it succeeded, but a same-container verifier can still inherit a planted pytest plugin, conftest.py, or executable. Run Python in isolated mode where practical and grade in a separate verifier container when tampering matters; the separate container keeps the agent's filesystem changes away from the verifier.

And read the passing runs, not only the failing ones, because that's where a hack shows up. Harbor's harbor analyze command reviews trajectories for this, looking for "modifications to test files," writes to the reward file, and access to the solution directory.

Flaky tasks

A flaky task is one whose result changes when nothing about the agent's work has changed. It's worse than a hard task, because it teaches everyone to ignore failures. The running example has an obvious source of flakiness: the shipper runs in the background, so the test has to wait for it. A fixed time.sleep(1) fails on a slow machine and wastes time on a fast one; the test above instead polls until the output has 500 lines or 20 seconds pass. Its one fixed wait comes after that, and only gives late duplicates time to appear.

Timing is only one source. The Terminal-Bench maintainers documented the others while fixing their own tasks:

Terminal-Bench 2.0 → 2.1 · 28 of 89 tasks fixed
cause
External dependencies
what went wrong
Nine tasks where external dependencies changed after the benchmark was built. The 2.0 post describes a YouTube download task that broke because of anti-bot changes: “a solution that worked one day might not work the next”
cause
Resources
what went wrong
Eight tasks had budgets too small for at least one valid approach, including the oracle or solutions from frontier models, to finish consistently
cause
Misspecification
what went wrong
Instructions and tests that disagreed, such as PostgreSQL in the instruction and Spark SQL in the tests

The machine matters too. Anthropic measured "the gap between the most- and least-resourced setups on Terminal-Bench 2.0" at 6 percentage points, and concluded that "leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched." The practical rules follow from all of this: pin the Python packages the tests use, avoid the network where the task allows it, set resources per task instead of relying on defaults, and run the reference solution many times. If the oracle doesn't pass every time, the task is flaky, whatever any agent scores.

One run is not a measurement

Agents are not deterministic. The same agent on the same task can pass on Monday and fail on Tuesday, so a single trial per task measures luck as much as ability. Two numbers describe several trials, and they answer different questions.

pass@k is the chance that at least one of k attempts succeeds: useful when a person will pick the best of several tries. pass^k is the chance that all k attempts succeed: what matters when the agent runs unattended and has to be right every time. The τ-bench paper, which introduced pass^k, reports that GPT-4o on its τ-retail domain had "> 60% average task success" but dropped below 25% on pass^8. For an agent that succeeds 75% of the time, assuming independent attempts, the two numbers move apart like this:

The same agent, measured two ways (75% success per attempt)

pass@k: at least one of k succeeds

Rises towards 100%: good for best-of-k with a human choosing.

pass^k: all k succeed

Falls towards 0%: what an unattended agent actually has to meet.

y: %

Computed as 1 − 0.25ᵏ and 0.75ᵏ. At k = 3, the agent almost always succeeds at least once, and succeeds every time only 42% of the time.

Harbor runs several attempts per task with -k and reports pass@k when rewards are pass or fail; it doesn't compute pass^k, so that has to come from the trial results. Terminal-Bench runs every agent "at least five times." And the extremes are informative on their own. Anthropic's guide puts it plainly: "a 0% pass rate across many trials (i.e. 0% pass@100) is most often a signal of a broken task, not an incapable agent."

Reading the runs

Scores tell you that something went wrong. Only the trajectory tells you what, and whether the fault is the agent's or the task's. Here is an illustrative failed trial of the running example:

One failed trial of fix-log-shipper (illustrative)

  1. 00:00

    Reads the code

    Opens /app/shipper.py and sees that the log is opened once, outside the loop.

  2. 00:41

    Reproduces the bug

    Starts the shipper, writes lines, renames the log by hand, and sees that nothing after the rename is shipped. A good sign: it checked before changing anything.

    mv app.log app.log.1 && touch app.log

  3. 01:58

    Checks too early

    Before trying one final read, it compares the path with the open handle. If rotation happened after its previous read, it switches immediately and abandons an unread tail in the old file.

    if rotated(): f = open(LOG) · else: line = f.readline()

  4. 02:30

    Tests its own fix and stops

    Rotates once by hand, with a pause between writing and rotating, and sees every line arrive. Declares the task done.

  5. verifier

    Abandons the unread tail

    The verifier freezes the reader, appends three lines to the old file, renames it, and resumes the process. The early identity check sees the new path first and switches before reading those lines.

    exact-list assertion fails · reward 0

    the agent's bug, not the task's

The question to ask of every failure is whose fault it is. Here the instruction said "exactly once … across any number of rotations," and the agent's own testing was too gentle to find the race, so the failure is fair. If every trial of every agent failed at the same line with the same error, the conclusion would be different; the debug tool described in appendix B.4 of the Terminal-Bench paper treats that pattern as evidence of a systematic instruction problem rather than agent variability. harbor view ./jobs shows each trajectory next to its verifier output, which makes the comparison quick. The same reading applies to passing runs, since that's where reward hacking hides.

Reviewing someone else's task

A review of a task is closer to a code review than a proofread. The reviewer's job is to find the ways the task would give the wrong score, which means running it as well as reading it. Anthropic's test for whether a task is well specified is a good one to hold every task to: "a good task is one where two domain experts would independently reach the same pass/fail verdict."

Instruction and tests agree

  • Every behaviour the tests check is in the instruction
  • Every requirement in the instruction is tested
  • Paths are absolute; output formats are defined
  • The tests accept every correct approach you can think of

Run it

  • The oracle scores 1, many times in a row
  • Doing nothing scores 0
  • Try to cheat it yourself, or run an agent told to
  • Watch at least one real agent attempt end to end

Nothing leaks

  • Tests and solution are not in the image
  • The solution isn't hard-coded against the tests
  • No answer is reachable over the network during the run
  • The verifier uses isolated Python or a separate container when tampering matters

Stable

  • Test dependencies pinned
  • Waits poll with a deadline instead of sleeping
  • Resources and timeouts fit every valid approach
  • No external service the task doesn't control

Rubrics for the work you can't test

The running example has a second half that no test can grade. Along with the fix, the agent writes a pull request description explaining what was wrong and how it was fixed. Whether that explanation is right is a judgment, so it's graded against a rubric: a list of criteria, each worth points.

The public work on rubric design agrees on a few rules. OpenAI's HealthBench, built with 262 physicians, gives each criterion a nonzero integer point value between −10 and 10, using negative points for undesirable behavior, and grades each criterion separately. Scale AI's "Rubrics as Rewards" paper asks for "7–20 self-contained items," each "independently actionable," and calls the negative ones pitfalls: they "help identify frequent or high-risk errors." Applied to the pull request description:

rubric · pull request description for fix-log-shipperpoints in HealthBench style; maximum 9
criterion
Names the cause: the shipper kept reading the renamed file and never opened the new one
kind
essential
points
+3
criterion
Explains how the fix detects rotation, by inode or an equivalent method
kind
essential
points
+3
criterion
Mentions the lines written just before the rename, and how the fix keeps them
kind
important
points
+2
criterion
Says how the fix was tested
kind
important
points
+1
criterion
Claims the fix also handles copy-and-truncate rotation when it doesn't
kind
pitfall
points
−3
criterion
Recommends restarting the shipper on every rotation
kind
pitfall
points
−2

Each criterion is a yes or no, not a score out of five. "Did it name the cause" can be checked; "rate the quality of the explanation" produces numbers that cluster in the middle and mean nothing. The second criterion says "or an equivalent method," for the same reason the tests accept both fixes: a rubric is also a test, and it can be narrow too. Graded against one description that explains the cause and the inode check and how it was tested, misses the last-lines race, and wrongly claims to handle copy-and-truncate, the score is (3 + 3 + 1 − 3) ÷ 9 = 4 ÷ 9, or 0.44.

Most rubric grading at volume is done by a model, the "LLM judge," and judges have measured biases. In one of the first large studies of LLM judges, researchers found "position, verbosity, and self-enhancement biases": GPT-4 favoured its own answers "with a 10% higher win rate," and only GPT-4 gave consistent verdicts more than 60% of the time when the order of two answers was swapped. The authors caution that the higher self-win rate is correlation and does not by itself prove self-preference. A separate study found that Vicuna-13B "could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator" just by changing the order the answers appeared in. The mitigations are the same ones careful teams use: grade one criterion per call, never compare two answers without also swapping them, judge with a different model family from the one being graded, and measure the judge against human labels. Mercor's own public benchmark paper, APEX-Agents, reports doing the last one: its judge agreed with 747 human labels 98.5% of the time.

From benchmark to production

Everything above is about tasks built to measure an agent. A team shipping an agent needs the same ideas arranged as a loop around its product:

  1. Test the tools. If a tool is broken, the agent looks incompetent and every eval above it measures the wrong thing. These tests need no model at all.
  2. Keep a set of verified tasks from your own product, built like the running example, and run it on every change that affects the agent. Real failures from production make the best tasks: freeze the situation, write the task, add the check.
  3. Grade what can't be tested with rubrics and an audited judge, on a sample large enough to show a change.
  4. Sample live traffic into human review, because real users do things no task set anticipated.
  1. Changeprompt, model, or tool
  2. Task setk trials each; gate on pass^k
  3. Shadow runsreal inputs, no side effects
  4. Judge and human samplerubric deltas vs baseline
  5. Ship or roll back
A release loop for an agent: verified tasks in CI, shadow runs on real inputs, rubric deltas, a human spot-check, then the decision.

The loop only works if the tasks in it deserve trust, which brings the chapter back to where it started. A benchmark as widely used as SWE-bench Verified still carried unfair tests when it was re-audited about 18 months after launch. The same will be true of any task set that nobody runs the oracle against, nobody tries to cheat, and nobody reads the runs of.

A checklist

Writing a task

  • The outcome is checked by running something, where possible
  • Instruction and tests describe the same behaviour
  • Tests accept every correct method and reject extra or duplicated output
  • Tests and solution stay out of the agent's container

Before trusting it

  • Oracle scores 1 on every run
  • Doing nothing scores 0
  • A cheating agent scores 0
  • Someone other than the author has run it

Running agents

  • Several trials per task; report pass^k where reliability matters
  • Resources and timeouts documented with the scores
  • Passing and failing runs both read
  • A 0% task investigated as a possibly broken task

Rubrics and judges

  • Yes/no criteria, one failure mode each
  • Pitfalls with negative points
  • One criterion per judge call; answer order swapped
  • Judge agreement with human labels measured

Sources