In August 2024, OpenAI had 93 software developers review 1,699 tasks from SWE-bench, one of the most widely used benchmarks for coding agents. About 539 survived the filtering process, and 500 were sampled for SWE-bench Verified. Some tasks described the problem too vaguely to solve; many more had tests that "may unfairly mark valid solutions as incorrect." In February 2026, OpenAI audited the Verified tasks that its o3 model kept failing and found that at least 59.4% had flawed tests that "reject functionally correct submissions."
So when someone says a new agent scores 62%, there are two questions behind the number. Do the tasks measure what they claim to measure? And would the agent score 62% again tomorrow? Most of the work in evaluating agents is making the answer to both questions yes, and it happens long before anyone looks at a leaderboard.
At Mercor I work on this part-time. I write agent tasks in Harbor, an open-source framework for running and grading agents; run agents on them and read what they did; review tasks other authors wrote; and write rubrics for work that can't be checked by a test. None of that work appears here. The task this chapter follows is one I made up for it, and every number comes from a public source.
Five rules shape the rest of this chapter:
- Where the outcome can be checked by running something, grade it with a test. A test is cheaper and faster than a judge, and much harder to argue with.
- Don't trust a task until three checks pass: the reference solution scores full marks, doing nothing scores zero, and an agent told to cheat scores zero.
- Run every task several times, and report how often the agent succeeds every time, not only whether it ever did.
- Read the runs. A score says that something happened; the run says what.
- Use rubrics only for what can't be tested, and write each criterion so that two experts would give the same answer.
How to read the diagrams
- The task: instruction, environment, files
- The agent and what it did
- The verifier and passing results
- Reference solutions and task checks
- Failures: cheats, flaky or unfair tests
The words this chapter uses
Evaluation has its own vocabulary, and most of the confusion in discussions about agents comes from mixing these terms up. The running example comes in the next section; this table shows each term as it applies there.
- term
- agent
- meaning
- A model in a loop that reads, runs commands, edits files, and decides its next step
- in the running example
- A coding agent working inside a container
- term
- task
- meaning
- One problem to solve: an instruction, an environment to solve it in, and a way to check the result
- in the running example
- Fix a log shipper that loses lines when its log file rotates
- term
- trial (run)
- meaning
- One agent's attempt at one task
- in the running example
- One attempt, start to finish, in a fresh container
- term
- trajectory
- meaning
- The record of a trial: every command, file read, edit, and message
- in the running example
- The agent read shipper.py, rotated the log by hand, then edited the loop
- term
- verifier
- meaning
- The program that checks the result after the agent stops
- in the running example
- A test that rotates the log mid-stream and compares the shipped lines
- term
- reward
- meaning
- The number the verifier writes: usually 1 for pass, 0 for fail
- in the running example
- 1 if all 500 lines arrived exactly once and in order
- term
- reference solution
- meaning
- A known-good answer, used to prove the task can be solved. Harbor runs it with an agent called oracle
- in the running example
- A shell script that installs a correct shipper
- term
- rubric
- meaning
- A list of criteria for grading output that no test can check, such as an explanation
- in the running example
- Criteria for the pull request description
Two kinds of task
Some outcomes can be checked by running something: the service starts, the numbers match, the file contains exactly the right lines. Others can't: whether an explanation is correct and complete, whether a code review caught the real problem. The two need different graders, and they fail in different ways.
Graded by a test
Verifiable tasks
- Grader
- A script that inspects the final state
- Cost per grade
- Seconds of compute
- Main risk
- Tests that reject correct answers or accept cheats
- Example
- Every log line arrives exactly once after rotation
Graded against criteria
Rubric-graded tasks
- Grader
- An expert, or a model checked against experts
- Cost per grade
- A model call per criterion, or an expert's time
- Main risk
- Criteria that two graders read differently
- Example
- The pull request explains the real cause
Terminal-Bench, a benchmark of hard tasks that agents solve in a Linux terminal, is built entirely from the first kind. Its paper says the tests "do not test the agent's commands or console output. This is intentional as Terminal-Bench is an outcome-driven framework." Anthropic's guide to agent evals makes the same point from the other direction: "it's often better to grade what the agent produced, not the path it took." Where both kinds apply, as they do in the running example, the test decides whether the work is correct, and the rubric grades what the test can't see.
The running example: a log shipper that loses lines
A small service, shipper.py, follows an application's log file and copies every new line to an output file, standing in for a real log pipeline. Every night the log is rotated the way logrotate does it by default: the current file is renamed to app.log.1, and a new, empty app.log is created.
The shipper opens the log once and reads it forever. It never notices the rename:
How the shipper loses lines at rotation
- 02:59:58
Reading app.log normally
The shipper holds an open file handle. On Linux, the handle points at the file itself, which the filesystem identifies by an inode number, not at the name.
app.log → inode 4711 · shipped: lines 1–100
- 03:00:00
logrotate renames and recreates
app.log becomes app.log.1, still inode 4711. A new, empty app.log is created with inode 4712.
app.log.1 → 4711 · app.log → 4712
- 03:00:01
The application keeps writing
The application opens the log by name, so lines 101 onward go into the new file.
inode 4712: lines 101, 102, 103 …
- 03:00:01 →
The shipper waits on a file nobody writes to
Its handle still points at inode 4711. It reaches the end of that file and waits for more lines forever. Nothing after line 100 is ever shipped.
shipped: lines 1–100 · lost: everything after
A correct fix has two parts. At the end of the file, the shipper checks whether the name app.log now points at a different inode; if it does, the old file has been rotated away, so it opens the new one and reads it from the start. And before it lets go of the old file, it reads it to the end one more time, because the application may have written a few lines to it between the shipper's last read and the rename. It's easy to get the first part right and miss the second, and that gap is what the task is really testing.
Anatomy of a Harbor task
Harbor is an open-source framework "from the creators of Terminal-Bench for evaluating and optimizing agents and language models," and "the official harness for Terminal-Bench-2.0." Its documentation puts a task in one sentence: "A Harbor task is an instruction, environment, and test script." On disk, the running example looks like this:
fix-log-shipper/
├── instruction.md what the agent is asked to do
├── task.toml timeouts, resources, free-form metadata
├── environment/
│ ├── Dockerfile the container the agent works in
│ └── shipper.py the buggy service
├── solution/
│ └── solve.sh the reference solution (run by the oracle agent)
└── tests/
├── test.sh the verifier: writes the reward
└── test_rotation.pyEach file has a job, and most task bugs come from one of them doing another's job.
The instruction
The log shipper at /app/shipper.py follows /var/log/app/app.log and appends
every line to /logs/artifacts/shipped.log. When the log is rotated, it loses lines.
Fix /app/shipper.py so that every line written to /var/log/app/app.log reaches
/logs/artifacts/shipped.log exactly once and in order, across any number of rotations.
Rotation works like logrotate's default: the current file is renamed to
app.log.1 and a new, empty app.log is created. Keep the output format: one
log line per output line, unchanged. The verifier performs rotations at least
one second apart and may write unread lines immediately before each rename.Every behaviour the tests will check is stated here: exactly once, in order, any number of rotations, the rotation method, the output format. Harbor's guidance is to "use absolute environment paths (/app/out.json not out.json) and define schemas for the output so instructions and tests agree." One detail surprises new authors: "the instruction is not passed to the agent as a file in the environment." The agent receives it as text, so it can't be found by searching the container.
The configuration
[metadata] # free-form: Harbor doesn't enforce these fields
author_name = "Robin Singh"
difficulty = "medium"
category = "debugging"
[agent]
timeout_sec = 900 # without this, Harbor enforces no agent timeout
[verifier]
timeout_sec = 600
[environment]
build_timeout_sec = 600
cpus = 1
memory_mb = 2048
network_mode = "no-network"The metadata table is, in the documentation's words, "arbitrary metadata provided by the task author"; difficulty and category are conventions, not schema. The timeouts and resources matter more than they look. Terminal-Bench's own maintainers wrote that "timeouts affect task difficulty," and later fixed eight tasks whose resource budgets were too small for at least one valid approach to finish reliably.
The environment
FROM python:3.12-slim
WORKDIR /app
COPY shipper.py /app/shipper.py
RUN mkdir -p /var/log/app /logs/artifactsWhat's missing matters as much as what's there. The tests and the reference solution are not copied into the image. Harbor's own task checker asks exactly this: "Are the tests/ folder or solution/ folder copied to the image? They should not be."
The verifier
#!/bin/bash
# Runs after the agent has stopped, in the same container.
mkdir -p /logs/verifier
pip install --quiet pytest==8.4.1
if pytest -q /tests/test_rotation.py; then
echo 1 > /logs/verifier/reward.txt
else
echo 0 > /logs/verifier/reward.txt
fiimport os, signal, subprocess, time
from pathlib import Path
LOG = Path("/var/log/app/app.log")
OUT = Path("/logs/artifacts/shipped.log")
def rotate():
os.rename(LOG, LOG.with_name("app.log.1")) # logrotate's default: rename…
LOG.touch() # …then create a new, empty file
def wait_for_lines(n, deadline=20.0):
end = time.monotonic() + deadline
while time.monotonic() < end:
if OUT.exists() and len(OUT.read_text().splitlines()) >= n:
return
time.sleep(0.05)
raise AssertionError(f"timed out waiting for {n} shipped lines")
def append_lines(first, last):
with LOG.open("a") as f:
for i in range(first, last + 1):
f.write(f"line {i}\n")
f.flush()
os.fsync(f.fileno())
def test_every_line_arrives_exactly_once_in_order():
OUT.unlink(missing_ok=True)
LOG.write_text("")
shipper = subprocess.Popen(["python3", "/app/shipper.py"])
try:
start = 1
for boundary in (100, 250, 400):
append_lines(start, boundary - 3)
wait_for_lines(boundary - 3)
# Freeze the reader, append an unread tail to the old file,
# rotate, then resume. A correct shipper drains the old handle.
os.kill(shipper.pid, signal.SIGSTOP)
append_lines(boundary - 2, boundary)
rotate()
os.kill(shipper.pid, signal.SIGCONT)
wait_for_lines(boundary)
time.sleep(1.0)
start = boundary + 1
append_lines(start, 500)
wait_for_lines(500)
time.sleep(0.5) # give late duplicates a chance to show up
assert OUT.read_text().splitlines() == [f"line {i}" for i in range(1, 501)]
finally:
shipper.kill()The test starts the shipper, writes 500 lines, and controls the race around
three rotations. SIGSTOP freezes the process; three unread lines are appended
to the old file, the file is renamed, and SIGCONT resumes the reader. The
wait_for_lines handshake proves each rotation completed before the next one.
Comparing the complete list catches loss, duplication, reordering, and altered
text. The script then writes Harbor's numerical reward under
/logs/verifier/.
The reference solution
#!/bin/bash
cat > /app/shipper.py <<'EOF'
import os, time
from pathlib import Path
LOG, OUT = "/var/log/app/app.log", "/logs/artifacts/shipped.log"
def open_log():
while True:
try:
f = open(LOG)
return f, os.fstat(f.fileno())
except FileNotFoundError:
time.sleep(0.05)
Path(OUT).parent.mkdir(parents=True, exist_ok=True)
f, opened_stat = open_log()
with open(OUT, "a") as out:
while True:
line = f.readline()
if line:
out.write(line)
out.flush()
continue
try:
rotated = not os.path.samestat(os.stat(LOG), opened_stat)
except FileNotFoundError:
rotated = False
if rotated:
for line in f: # lines written just before the rename
out.write(line)
out.flush()
f.close()
f, opened_stat = open_log() # the new file, from its first line
else:
time.sleep(0.05)
EOFThe solution folder is, per the documentation, "a reference script used by the oracle agent to sanity-check that a task is solvable." It's optional for published benchmarks, but "without solution/, the Oracle agent cannot run," and without the oracle you have no proof that your own task can be passed.
What happens when you run it
A run in Harbor is called a trial: "one agent's attempt at completing one task." A job is a collection of trials, run in parallel. The order of events inside one trial explains most of the design above:
One trial, from command to reward
harbor run
The command names the task, the agent, the model, and how many attempts to make.
harbor run -p ./fix-log-shipper -a claude-code -m <model> -k 5
The environment is built
The container is built from environment/, within the build timeout.
Dockerfile → image · build_timeout_sec 600
The agent works
The agent receives the instruction as text and works in the container until it stops or hits its timeout. The tests and the solution are not in the container.
agent/trajectory.json records every step
The tests are uploaded
Only now does Harbor copy tests/ into the container.
tests/ → /tests
The verifier runs
test.sh runs within its own timeout and writes the reward.
/logs/verifier/reward.txt → 1 or 0
The reward is collected
Harbor reads the reward file. A missing or non-numeric reward is an error, not a zero, and by default isn't retried.
result.json · /logs/artifacts downloaded · container deleted
The job is summarised
Rewards are averaged across tasks; missing rewards count as 0. With binary rewards and several attempts, Harbor also reports pass@k.
harbor view ./jobs to read the trials
The detail that matters most is the fourth step. In the documentation's words, "the tests/ directory is uploaded to the environment at /tests/ after the agent runs." An agent can't read tests that don't exist yet. Harbor can go further and run the verifier in a separate container, where "the agent's filesystem changes are not inherited." Files deliberately published under /logs/artifacts/, plus any paths declared as artifacts, are copied into that verifier. The running example puts shipped.log there so the isolated verifier can inspect the outcome without inheriting the rest of the agent's filesystem.
Three checks before anyone trusts a task
A task is software, and like software it has bugs. Terminal-Bench's continuous-integration checks for every task are a good minimum, and all three can run before any real agent touches the task:
- check
- oracle
- what it proves
- The task can be solved, and the tests accept a correct answer
- command
- harbor run -p ./fix-log-shipper -a oracle
- must score
- 1
- check
- nop (does nothing)
- what it proves
- The tests don't pass the starting state: the bug is real and the tests catch it
- command
- harbor run -p ./fix-log-shipper -a nop
- must score
- 0
- check
- cheating agent
- what it proves
- An agent told to game the grader can't: no reading answers, no editing the checker
- command
- an agent prompted to reward-hack
- must score
- 0
The Terminal-Bench paper describes the same gate for its 2.0 release: "an automated workflow ran the task's oracle solution to ensure solvability," and "a no-op 'dummy' agent should fail the task." Each of its 89 tasks was then "manually verified by three human reviewers for correctness," using about three hours of combined reviewer attention on average. Harbor packages a first pass of that review as a command: harbor check <task-dir> grades a task against a rubric of common problems, including whether the tested behaviour is described in the instruction, whether Python dependencies are pinned, and whether the solution is hard-coded.
Tests that grade the result, not the method
A test can fail a task in two directions. It can reject a correct answer, or it can accept a wrong one. OpenAI's 2026 audit of SWE-bench Verified sorted the flawed tests it found into "narrow" tests, which reject valid solutions, and "wide" tests, which check for behaviour the task never asked for.
- review
- August 2024: 1,699 original tasks, each labelled by three developers
- finding
- 38.3% flagged for underspecified problem statements; 61.1% for unit tests that may unfairly mark valid solutions as incorrect; 68.3% filtered out, leaving 500
- review
- February 2026: 138 Verified tasks that OpenAI's o3 didn't solve consistently over 64 runs, each reviewed by at least six engineers
- finding
- 59.4% had flawed tests that reject correct submissions: 35.5% narrow, 18.8% wide, 5.1% other
- review
- The narrow-test example
- finding
- A test that imports a function called get_annotation, a name the problem description never mentions. Any fix that names the function differently fails
The same thing is easy to do in the running example. Suppose a task author, trying to be thorough, adds a test that the fix reads st_ino directly. That is one way to recognise a rename, but it is not the only correct one. Python's os.path.samestat compares the device and inode from two stat results; an inotify watcher could also report the move. A test tied to one implementation rejects the others:
| fix | narrow test: “uses st_ino” | outcome test: “500 lines, once, in order” |
|---|---|---|
| Fix A: compares inode numbers | ||
| checks os.stat(LOG).st_ino | pass | pass |
| Fix B: compares complete file identities | ||
| uses os.path.samestat on the open handle and current path | fail: wrongly | pass |
Terminal-Bench's criterion for this is the cleanest statement I know: "the unit tests will pass if and only if the container ends in an acceptable state." Its guidance for task authors says the same in practical terms: "Tests should validate outcomes rather than implementation details. They should allow alternate correct solutions, avoid brittle source-code checks, and protect against reward hacking or accidental leakage of the reference answer."
Wide tests are the mirror image, and they usually come from the instruction and the tests drifting apart. When the Terminal-Bench maintainers fixed 28 of their 89 tasks for version 2.1, one example was a task whose "tests expected Spark SQL output, while the instructions asked for PostgreSQL." The fix is procedural: every behaviour the tests check must be written in the instruction, and every requirement in the instruction must be checked by a test.
When the agent games the test
Agents trained with reinforcement learning are rewarded for passing checks, and some learn to pass the check instead of doing the task. This is called reward hacking, and it's well documented:
Published cases of agents gaming their graders
- Feb 2025
Anthropic: special-casing tests
The Claude 3.7 Sonnet system card: the model “occasionally resorts to special-casing in order to pass test cases,” most often “directly returning expected test values rather than implementing general solutions, but also includes modifying the problematic tests themselves.” Anthropic's later Sonnet 4.5 system card reported much lower rates on updated reward-hacking evaluations, while noting that hard-coding and special-casing still occurred.
Claude 3.7 and Sonnet 4.5 system cards
- Mar 2025
OpenAI: “hack verify to always return true”
Reasoning traces from coding tasks included plans like “Test only calls analyze.prove and verify. They don't inspect details. So we can hack verify to always return true.” The accompanying paper lists exiting early with exit(0), raising SkipTest, and writing stubs where test coverage is thin.
OpenAI, Mar 2025; arXiv 2503.11926
- Jun 2025
METR: patching the scorer
o3 “decides to patch the competition evaluation function so that it judges every submission as successful.” On one task it hacked in every run. Telling the model not to cheat “had a nearly negligible effect.”
METR, Recent Frontier Models Are Reward Hacking
- Jan 2026
Terminal-Bench: an agent built to find exploits
The benchmark's authors ran an adversarial agent against their own tasks. It found “monkey-patching test environments, guessing answers, and generating all possible answers with the hope that tests might only check existence of a correct answer but not absence of wrong answers.”
arXiv 2601.11868, appendix B.4
- Apr 2026
Terminal-Bench leaderboard submissions
The integrity review found different failures: OB-1 had modified timeouts in its original Terminal-Bench 1.0 submission; Pilot uploaded the tests folder; and ForgeCode runs downloaded solutions into AGENTS.md. The maintainers said reward-hacked trials would score 0 and announced an agent judge for passing trials.
tbench.ai, Leaderboard Integrity Update
Anthropic notes that its own measurements come from evals "explicitly designed to stress-test hacking propensities" and don't reflect real-world rates, so these cases say that hacking happens, not how often. For a task author, that's enough: any gap in the grading will eventually be found.
The Terminal-Bench exploit list maps directly onto the running example. A shipper that simply writes every line twice contains all 500 expected lines. A test that only checks that each expected line appears somewhere would pass it:
| output | weak test: “every line appears” | strict test: “exactly this list” |
|---|---|---|
| Correct shipper | ||
| 500 lines, once each, in order | pass | pass |
| Shipper that ships every line twice | ||
| 1,000 lines: each expected line, twice | pass: wrongly | fail |
The defences are mostly about where things live and what the test trusts. Keep answers out of reach: Harbor uploads the tests only after the agent finishes, the solution never goes into the image, and a task can cut the agent off from the network during its run, which blocks the most common loophole on the Terminal-Bench leaderboard. Check for wrong answers as well as right ones, by comparing the exact output, so duplicates and extra lines fail. This test writes its own log lines and starts the shipper itself instead of asking the agent's code whether it succeeded, but a same-container verifier can still inherit a planted pytest plugin, conftest.py, or executable. Run Python in isolated mode where practical and grade in a separate verifier container when tampering matters; the separate container keeps the agent's filesystem changes away from the verifier.
And read the passing runs, not only the failing ones, because that's where a hack shows up. Harbor's harbor analyze command reviews trajectories for this, looking for "modifications to test files," writes to the reward file, and access to the solution directory.
Flaky tasks
A flaky task is one whose result changes when nothing about the agent's work has changed. It's worse than a hard task, because it teaches everyone to ignore failures. The running example has an obvious source of flakiness: the shipper runs in the background, so the test has to wait for it. A fixed time.sleep(1) fails on a slow machine and wastes time on a fast one; the test above instead polls until the output has 500 lines or 20 seconds pass. Its one fixed wait comes after that, and only gives late duplicates time to appear.
Timing is only one source. The Terminal-Bench maintainers documented the others while fixing their own tasks:
- cause
- External dependencies
- what went wrong
- Nine tasks where external dependencies changed after the benchmark was built. The 2.0 post describes a YouTube download task that broke because of anti-bot changes: “a solution that worked one day might not work the next”
- cause
- Resources
- what went wrong
- Eight tasks had budgets too small for at least one valid approach, including the oracle or solutions from frontier models, to finish consistently
- cause
- Misspecification
- what went wrong
- Instructions and tests that disagreed, such as PostgreSQL in the instruction and Spark SQL in the tests
The machine matters too. Anthropic measured "the gap between the most- and least-resourced setups on Terminal-Bench 2.0" at 6 percentage points, and concluded that "leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched." The practical rules follow from all of this: pin the Python packages the tests use, avoid the network where the task allows it, set resources per task instead of relying on defaults, and run the reference solution many times. If the oracle doesn't pass every time, the task is flaky, whatever any agent scores.
One run is not a measurement
Agents are not deterministic. The same agent on the same task can pass on Monday and fail on Tuesday, so a single trial per task measures luck as much as ability. Two numbers describe several trials, and they answer different questions.
pass@k is the chance that at least one of k attempts succeeds: useful when a person will pick the best of several tries. pass^k is the chance that all k attempts succeed: what matters when the agent runs unattended and has to be right every time. The τ-bench paper, which introduced pass^k, reports that GPT-4o on its τ-retail domain had "> 60% average task success" but dropped below 25% on pass^8. For an agent that succeeds 75% of the time, assuming independent attempts, the two numbers move apart like this:
The same agent, measured two ways (75% success per attempt)
pass@k: at least one of k succeeds
Rises towards 100%: good for best-of-k with a human choosing.
pass^k: all k succeed
Falls towards 0%: what an unattended agent actually has to meet.
y: %
Harbor runs several attempts per task with -k and reports pass@k when rewards are pass or fail; it doesn't compute pass^k, so that has to come from the trial results. Terminal-Bench runs every agent "at least five times." And the extremes are informative on their own. Anthropic's guide puts it plainly: "a 0% pass rate across many trials (i.e. 0% pass@100) is most often a signal of a broken task, not an incapable agent."
Reading the runs
Scores tell you that something went wrong. Only the trajectory tells you what, and whether the fault is the agent's or the task's. Here is an illustrative failed trial of the running example:
One failed trial of fix-log-shipper (illustrative)
- 00:00
Reads the code
Opens /app/shipper.py and sees that the log is opened once, outside the loop.
- 00:41
Reproduces the bug
Starts the shipper, writes lines, renames the log by hand, and sees that nothing after the rename is shipped. A good sign: it checked before changing anything.
mv app.log app.log.1 && touch app.log
- 01:58
Checks too early
Before trying one final read, it compares the path with the open handle. If rotation happened after its previous read, it switches immediately and abandons an unread tail in the old file.
if rotated(): f = open(LOG) · else: line = f.readline()
- 02:30
Tests its own fix and stops
Rotates once by hand, with a pause between writing and rotating, and sees every line arrive. Declares the task done.
- verifier
Abandons the unread tail
The verifier freezes the reader, appends three lines to the old file, renames it, and resumes the process. The early identity check sees the new path first and switches before reading those lines.
exact-list assertion fails · reward 0
the agent's bug, not the task's
The question to ask of every failure is whose fault it is. Here the instruction said "exactly once … across any number of rotations," and the agent's own testing was too gentle to find the race, so the failure is fair. If every trial of every agent failed at the same line with the same error, the conclusion would be different; the debug tool described in appendix B.4 of the Terminal-Bench paper treats that pattern as evidence of a systematic instruction problem rather than agent variability. harbor view ./jobs shows each trajectory next to its verifier output, which makes the comparison quick. The same reading applies to passing runs, since that's where reward hacking hides.
Reviewing someone else's task
A review of a task is closer to a code review than a proofread. The reviewer's job is to find the ways the task would give the wrong score, which means running it as well as reading it. Anthropic's test for whether a task is well specified is a good one to hold every task to: "a good task is one where two domain experts would independently reach the same pass/fail verdict."
Instruction and tests agree
- Every behaviour the tests check is in the instruction
- Every requirement in the instruction is tested
- Paths are absolute; output formats are defined
- The tests accept every correct approach you can think of
Run it
- The oracle scores 1, many times in a row
- Doing nothing scores 0
- Try to cheat it yourself, or run an agent told to
- Watch at least one real agent attempt end to end
Nothing leaks
- Tests and solution are not in the image
- The solution isn't hard-coded against the tests
- No answer is reachable over the network during the run
- The verifier uses isolated Python or a separate container when tampering matters
Stable
- Test dependencies pinned
- Waits poll with a deadline instead of sleeping
- Resources and timeouts fit every valid approach
- No external service the task doesn't control
Rubrics for the work you can't test
The running example has a second half that no test can grade. Along with the fix, the agent writes a pull request description explaining what was wrong and how it was fixed. Whether that explanation is right is a judgment, so it's graded against a rubric: a list of criteria, each worth points.
The public work on rubric design agrees on a few rules. OpenAI's HealthBench, built with 262 physicians, gives each criterion a nonzero integer point value between −10 and 10, using negative points for undesirable behavior, and grades each criterion separately. Scale AI's "Rubrics as Rewards" paper asks for "7–20 self-contained items," each "independently actionable," and calls the negative ones pitfalls: they "help identify frequent or high-risk errors." Applied to the pull request description:
- criterion
- Names the cause: the shipper kept reading the renamed file and never opened the new one
- kind
- essential
- points
- +3
- criterion
- Explains how the fix detects rotation, by inode or an equivalent method
- kind
- essential
- points
- +3
- criterion
- Mentions the lines written just before the rename, and how the fix keeps them
- kind
- important
- points
- +2
- criterion
- Says how the fix was tested
- kind
- important
- points
- +1
- criterion
- Claims the fix also handles copy-and-truncate rotation when it doesn't
- kind
- pitfall
- points
- −3
- criterion
- Recommends restarting the shipper on every rotation
- kind
- pitfall
- points
- −2
Each criterion is a yes or no, not a score out of five. "Did it name the cause" can be checked; "rate the quality of the explanation" produces numbers that cluster in the middle and mean nothing. The second criterion says "or an equivalent method," for the same reason the tests accept both fixes: a rubric is also a test, and it can be narrow too. Graded against one description that explains the cause and the inode check and how it was tested, misses the last-lines race, and wrongly claims to handle copy-and-truncate, the score is (3 + 3 + 1 − 3) ÷ 9 = 4 ÷ 9, or 0.44.
Most rubric grading at volume is done by a model, the "LLM judge," and judges have measured biases. In one of the first large studies of LLM judges, researchers found "position, verbosity, and self-enhancement biases": GPT-4 favoured its own answers "with a 10% higher win rate," and only GPT-4 gave consistent verdicts more than 60% of the time when the order of two answers was swapped. The authors caution that the higher self-win rate is correlation and does not by itself prove self-preference. A separate study found that Vicuna-13B "could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator" just by changing the order the answers appeared in. The mitigations are the same ones careful teams use: grade one criterion per call, never compare two answers without also swapping them, judge with a different model family from the one being graded, and measure the judge against human labels. Mercor's own public benchmark paper, APEX-Agents, reports doing the last one: its judge agreed with 747 human labels 98.5% of the time.
From benchmark to production
Everything above is about tasks built to measure an agent. A team shipping an agent needs the same ideas arranged as a loop around its product:
- Test the tools. If a tool is broken, the agent looks incompetent and every eval above it measures the wrong thing. These tests need no model at all.
- Keep a set of verified tasks from your own product, built like the running example, and run it on every change that affects the agent. Real failures from production make the best tasks: freeze the situation, write the task, add the check.
- Grade what can't be tested with rubrics and an audited judge, on a sample large enough to show a change.
- Sample live traffic into human review, because real users do things no task set anticipated.
- Changeprompt, model, or tool
- Task setk trials each; gate on pass^k
- Shadow runsreal inputs, no side effects
- Judge and human samplerubric deltas vs baseline
- Ship or roll back
The loop only works if the tasks in it deserve trust, which brings the chapter back to where it started. A benchmark as widely used as SWE-bench Verified still carried unfair tests when it was re-audited about 18 months after launch. The same will be true of any task set that nobody runs the oracle against, nobody tries to cheat, and nobody reads the runs of.
A checklist
Writing a task
- The outcome is checked by running something, where possible
- Instruction and tests describe the same behaviour
- Tests accept every correct method and reject extra or duplicated output
- Tests and solution stay out of the agent's container
Before trusting it
- Oracle scores 1 on every run
- Doing nothing scores 0
- A cheating agent scores 0
- Someone other than the author has run it
Running agents
- Several trials per task; report pass^k where reliability matters
- Resources and timeouts documented with the scores
- Passing and failing runs both read
- A 0% task investigated as a possibly broken task
Rubrics and judges
- Yes/no criteria, one failure mode each
- Pitfalls with negative points
- One criterion per judge call; answer order swapped
- Judge agreement with human labels measured
Sources
- Harbor: repository (Apache-2.0) and documentation on tasks, instructions, configuration, solutions, the verifier, separate verifiers, network policies, running jobs, and metrics
- Terminal-Bench: Merrill, Shaw et al., Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces (2026); tbench.ai posts on 2.0 and Harbor, what makes a good task, 2.1, continuous benchmarks, timeouts, and leaderboard integrity
- OpenAI: Introducing SWE-bench Verified (2024), Why SWE-bench Verified no longer measures frontier coding capabilities (2026), Detecting misbehavior in frontier reasoning models (2025) and Baker et al., Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (2025), Arora et al., HealthBench (2025)
- Anthropic: Claude 3.7 Sonnet System Card (2025), Claude Sonnet 4.5 System Card (2025), Demystifying evals for AI agents (2026), Quantifying infrastructure noise in agentic coding evals (2026)
- METR, Recent Frontier Models Are Reward Hacking (2025)
- Yao et al., τ-bench (2024); Chen et al., Evaluating Large Language Models Trained on Code (2021)
- Gunjal et al., Rubrics as Rewards (2025); Mercor, APEX-Agents (2026)
- Zheng et al., Judging LLM-as-a-Judge (2023); Wang et al., Large Language Models are not Fair Evaluators (2023)