Skip to main content
Once an evaluation request starts writing output, it stores task results, run records, aggregate metrics, and logs in one run directory. This page introduces that directory and helps you find the right file for what you want to inspect. The following pages explain each artifact type and how task results become aggregate metrics. If an evaluation fails preflight before the run directory is created, or if you use launch --dry-run, no result directory is generated.

Directory Layout

The command determines the run directory. For agentcompass run:
For agentcompass launch:
The combined component IDs and request names are normalized to safe directory components. The default result root is results. Each request name gives a launch request its own output namespace, so different requests can use the same run ID even when their Model and Benchmark match. A complete run typically generates the following directories and files inside its run directory:
If run-name is not set, that path segment is omitted. Task directories are grouped by the most severe final issue of any attempt: fatal/ if any attempt ends with a fatal issue, otherwise error/ if any attempt has an error, otherwise normal/, which also holds warning-only results. Unfinished tasks stay in running/ and move to their group when the task result is written; they are not counted as results. Each attempt owns its checkpoint.json; retries/ is created only when retry diagnostics or previous execution outputs need to be retained. Analysis summaries appear only when there are analysis results to aggregate. If a run stops during preflight, task execution, or summary generation, its directory may contain only the artifacts written up to that point.

Legacy

The previous layout stored each task’s complete result in one file and kept checkpoints, retry diagnostics, and logs in separate directories. The comparison below omits unchanged run-level files such as metrics and progress. <task-key> means the readable task ID followed by its full SHA-256 suffix. Previous layout:
Current layout:
Previously, a task file contained all its attempts; the _error_ filename variant indicated an execution or evaluation error. Now task.json stores shared task metadata and the mapping such as "1": "attempt-1". Each attempt owns its result, checkpoint, collected artifacts, and retry records. Issues are recorded in the result; the state directory groups tasks by their most severe final issue. The intermediate details/<task-key>/ layout without a state directory is still readable. Optional directories appear only when needed.

Reuse Support

  • Runs with a supported run schema, the same Benchmark ID, and the same attempt plan support --reuse, subject to task/attempt and artifact checks. Missing or historical fingerprint fields do not affect reuse.
  • Unsupported run schemas still require a new run. Renaming directories or moving files does not migrate a schema.
  • Runtime result loading, reuse, summary regeneration, and offline analysis use the current task.json and per-attempt layout, and also read details/<task-key>/ directories without a state directory. Legacy flat details and earlier indexed layouts are not supported or automatically upgraded.
  • Reuse always writes to a new run directory. Tasks copied as final results go to their state directory; tasks whose attempts will be retried are placed in running/ until they finish.

Where to Start

details/<state>/<task-key>/task.json stores shared task metadata and the attempt-ID mapping; each attempt owns its result.json. Readers combine these files and derive task retry totals without a separate task-level result file. metrics.json is the canonical run-level report, while summary.md is its readable presentation: it keeps the traditional metric layout for k=1 and shows the repeated-attempt plan and series for k>1. A later agentcompass summary invocation rereads the details and persisted attempt plan. When analysis is enabled, output for each evaluation attempt is stored under analysis_result and then aggregated separately. Progress files, logs, checkpoints, and attempt retries/ support monitoring, recovery, and diagnosis; retry diagnostics do not directly contribute to Benchmark metrics.

Data, Cache, and Output Directories

Benchmark data and evaluation results are stored in different directories. Use this table to choose the appropriate setting: In a configuration file, use runtime.data_dir and runtime.results_dir to set the root directories. For a single evaluation request, you can also pass the corresponding CLI options. Because run-name and run-id are output settings for an individual request, place them under that request’s output in a multi-evaluation orchestration file. See agentcompass run and agentcompass launch.