Skip to main content
GAIA (arxiv, “GAIA: a benchmark for General AI Assistants”) evaluates general AI assistants on real-world questions that require tool use, web browsing, and multi-step reasoning. For each question the assistant produces a short final answer, which an LLM judge then grades as correct or not against the ground truth. GAIA uses the GAIA 2023 validation split. Inference and judging both run in the local process (host_process). The harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it. Judging is single-sided: the judge only compares the agent-under-test’s answer against the ground truth, without comparing to any baseline.

How it works

A GAIA run has two stages — inference and judging.

Inference and judging

  • Inference. The model under test acts as a search agent and, driven by the harness (default naive_search_agent), completes multi-turn tool loops such as search / visit per task, producing a short natural-language final answer.
  • Judging. The judge model (judge_model) receives “question + ground truth + answer under test” and grades it with the built-in A/B/C protocol. The judge compares only the final answer, ignoring reasoning and formatting differences; equivalent expressions are accepted. The judge and the model under test are two separate endpoints; judge_model must be specified explicitly.

The A/B/C verdict

The judge returns exactly one verdict, and only A counts as correct:
  • A — CORRECT: the answer semantically matches the ground truth (equivalent expressions and formatting allowed).
  • B — INCORRECT: any deviation from the ground truth.
  • C — INCOMPLETE / REPETITIVE / REFUSAL: an invalid answer (cut off mid-sentence, looping repetition, or an explicit refusal).

Parameters

Pass a JSON object via --benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.

Parameter reference

ParameterTypeDefaultChoices / valuesDescription
judge_modeldictnull{id, base_url, api_key, api_protocol, params}Judge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*.
categorystring / list”all""all”, 1, 2, 3Filter by GAIA difficulty level (1 = easiest to 3 = hardest); “all” = no filter, a list takes the union. Task counts by level — 1 (39), 2 (52), 3 (12); 103 in total.
Shared parameters such as k, avgk, and sample_ids follow the conventions in Benchmark Parameters.

Judge model spec

judge_model is passed as a dict: {"id","base_url","api_key","api_protocol","params"}, pointing to the judge model’s own endpoint, with inference parameters under params. We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — the A/B/C criterion (semantic match) is relatively objective, so a mid-sized model suffices. AgentCompass recommends Qwen3.6-35B-A3B.

Run examples

The GAIA run command takes the form agentcompass run gaia <harness> <model>, whose three positional arguments are:
  • gaia — the benchmark id;
  • <harness> — the harness that drives the model under test through the search loop, defaulting to naive_search_agent; its own configuration is passed via --harness-params;
  • <model> — the model under test, i.e. the agent that performs retrieval and answering; its access credentials are passed via --model-base-url / --model-api-key.
Run configuration is split into two JSON blocks: --benchmark-params carries benchmark-level configuration (judge model, data filtering; see the Parameter reference above), and --harness-params carries the naive_search_agent harness’s own configuration (enabled tools, Serper / Jina keys, iterations, timeout, etc.; see the full list in NaiveSearchAgent harness). Both can also be written into the benchmark.params / harness.params blocks of --config, with the CLI winning on shared keys. In the examples below, --harness-params always passes the Serper and Jina keys required for retrieval directly via serper_api_key / jina_api_key (the search / visit tools of naive_search_agent depend on them); the examples differ only in --benchmark-params.
Use sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

Outputs

A run produces two kinds of results, both under results/gaia/<model>/<run>/: aggregate metrics (summary.md, overall performance) and per-task details (details/, per-task grading).

Aggregate metrics (summary.md)

summary.md summarizes the overall performance of the run, in two parts — a run overview and the metrics. Run overview Metrics There is a single headline metric, accuracy: the share of tasks judged correct. A task counts as correct (scored 1, otherwise 0) if and only if the judge returns verdict A; accuracy is the average over all tasks.

Per-task details (details/)

Each task has one JSON file, in which the judge’s grading for the task is recorded under the extra.scoring field: Only the parsed verdict is persisted here; the answer under test and ground truth are kept for tracing, while the full trajectory is written alongside in the same task file.