Skip to main content
Terminal-Bench 2.1 is the AgentCompass entry for the Terminal-Bench 2.1 task collection. It uses the same containerized execution and verification flow as Terminal-Bench 2, normally with the terminus2 harness.

How it works

  1. Load tasks. With the default dataset address, AgentCompass downloads Terminal-Bench 2.1 through the Harbor CLI. An explicitly configured regular Git source is shallow-cloned instead.
  2. Run the agent. The environment recipe selects the image declared by each task and starts the agent in its prepared terminal workspace.
  3. Verify the result. The task’s tests/test.sh is run through the Harbor verifier. A reward of 1 marks the task correct.

Parameters

Configure Terminal-Bench-specific options with --benchmark-params '{...}'.

Run examples

Use --benchmark-params for dataset and judge settings, --harness-params for agent and tool settings, and --execution-params for phase timeouts and multipliers. YAML uses benchmark.params, harness.params, and execution; explicit CLI values override YAML values. The recommended terminal agent is terminus2.
Verify the end-to-end flow works — sample_ids selects which case to run, with all other parameters using their defaults.

Other optional harnesses

codex and claude_code are two other harness options. Pass --recipe terminalbench2_1_docker_ac to use the AgentCompass prebuilt image. It includes download dependencies such as Node.js, npm, curl, and wget for Codex, Claude Code, and similar harnesses.
Omit --recipe to use the official task image. Because it does not include the Node bootstrap dependencies, provide the matching installation command explicitly.

Output

A run writes per-task details and the aggregate views summary.md and metrics.json under the run directory.

Aggregate metrics (summary.md)

The primary metric is binary correct: the Harbor verifier’s full reward (1) maps to true. At k=1, correct.native@1 is Terminal-Bench’s pass rate over evaluated observations. At k>1, the generic reducers can emit correct.avg@k and correct.pass@k, each with independent counts.

Per-task details (details/)

Each task JSON stores the binary observation at attempts.<N>.metrics.correct, together with execution status, the agent trajectory, Harness diagnostics, and raw verifier evidence. See Results.