> ## Documentation Index
> Fetch the complete documentation index at: https://agent-compass.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# BrowseComp-ZH

BrowseComp-ZH ([arxiv](https://arxiv.org/abs/2504.19314)) evaluates a browsing agent's web browsing ability in Chinese: given a question whose answer is a hard-to-find fact on the Chinese-language web — one that a single search cannot answer and that requires persistent, multi-step browsing to pin down — the agent produces a short final answer, which an **LLM judge** then grades as correct or not against the ground truth.

Like DeepSearchQA, BrowseComp-ZH uses single-sided judging. The judge only compares the agent-under-test's answer against the ground truth, without comparing to any baseline. Both inference and judging run in the local process (`host_process`) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

## How it works

A BrowseComp-ZH run has two stages — inference and judging.

### Inference and judging

* **Inference.** The model under test acts as a search agent and, driven by the harness (default [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent)), completes multi-turn tool loops such as search / visit per task, producing a short natural-language final answer.
* **Judging.** The judge model (`judge_model`) receives "question + ground truth + answer under test" and grades it with the built-in A/B/C protocol. The judge compares only the final answer, ignoring reasoning and formatting differences; equivalent expressions are accepted. The judge and the model under test are two separate endpoints; `judge_model` must be specified explicitly.

### The A/B/C verdict

The judge returns exactly one verdict, and only **A** counts as correct:

* **A — CORRECT:** the answer semantically matches the ground truth (equivalent expressions and formatting allowed).
* **B — INCORRECT:** any deviation from the ground truth.
* **C — INCOMPLETE / REPETITIVE / REFUSAL:** an invalid answer (cut off mid-sentence, looping repetition, or an explicit refusal).

## Parameters

Pass a JSON object via `--benchmark-params '{...}'`, or a `benchmark.params` block in the YAML given to `--config`; the CLI wins on shared keys. See the [Benchmark overview](/en/user_guide/modules/benchmarks/overview) for merge precedence.

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1040px', width:'100%'}}>
    <colgroup>
      <col width="18%" />

      <col width="16%" />

      <col width="15%" />

      <col width="20%" />

      <col width="31%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>judge\_model</code></td><td style={{whiteSpace:'nowrap'}}>dict</td><td style={{whiteSpace:'nowrap'}}><code>null</code></td><td><code>\{id, base\_url, api\_key, api\_protocol, params}</code></td><td>Judge model spec, <strong>required</strong> (see <a href="#judge-model-spec">Judge model spec</a>). It decides grading, and is not the CLI <code>--model-\*</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>category</code></td><td style={{whiteSpace:'nowrap'}}>string / list</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>"all"</code>, <code>影视</code>, <code>艺术</code>, <code>地理</code>, <code>音乐</code>, <code>历史</code>, <code>医学</code>, <code>电子游戏</code>, <code>科技</code>, <code>体育</code>, <code>政策法规</code>, <code>学术论文</code></td><td>Filter tasks by category; <code>"all"</code> = no filter, a list takes the union. Task counts by category — 影视 (45), 艺术 (40), 地理 (37), 音乐 (32), 历史 (29), 医学 (26), 电子游戏 (23), 科技 (22), 体育 (18), 政策法规 (10), 学术论文 (7); 289 in total.</td></tr>
    </tbody>
  </table>
</div>

Shared parameters such as `k`, `avgk`, and `sample_ids` follow the conventions in [Benchmark Parameters](/en/user_guide/modules/benchmarks/overview).

<a id="judge-model-spec" />

### Judge model spec

`judge_model` is passed as a dict: `{"id","base_url","api_key","api_protocol","params"}`, pointing to the judge model's own endpoint, with inference parameters under `params`.

We recommend **fixing a single judge** across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — the A/B/C criterion (semantic match) is relatively objective, so a mid-sized model suffices. AgentCompass recommends `Qwen3.6-35B-A3B`.

## Run examples

The BrowseComp-ZH run command takes the form `agentcompass run browsecomp_zh <harness> <model>`, whose three positional arguments are:

* `browsecomp_zh` — the benchmark id;
* `<harness>` — the harness that drives the model under test through the search loop, defaulting to [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent); its own configuration is passed via `--harness-params`;
* `<model>` — the model under test, i.e. the agent that performs retrieval and answering; its access credentials are passed via `--model-base-url` / `--model-api-key`.

Run configuration is split into two JSON blocks: `--benchmark-params` carries benchmark-level configuration (judge model, data filtering; see the [Parameter reference](#parameter-reference) above), and `--harness-params` carries the [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent) harness's own configuration (enabled tools, Serper / Jina keys, iterations, timeout, etc.; see the full list in [NaiveSearchAgent harness](/en/user_guide/modules/harnesses/naive_search_agent)). Both can also be written into the `benchmark.params` / `harness.params` blocks of `--config`, with the CLI winning on shared keys.

In the examples below, `--harness-params` always passes the Serper and Jina keys required for retrieval directly via `serper_api_key` / `jina_api_key` (the `search` / `visit` tools of `naive_search_agent` depend on them); the examples differ only in `--benchmark-params`.

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Use `sample_ids` to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

    ```bash theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "Qwen3.6-35B-A3B", "base_url": "https://your-judge-endpoint/v1", "api_key": "sk-…"},
        "sample_ids": ["1"]
      }' \
      --harness-params '{
        "serper_api_key": "your-serper-key",
        "jina_api_key": "your-jina-key"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom parameters">
    Evaluate only a subset of categories to focus analysis on a specific domain; also demonstrates narrowing the toolset and iterations in `--harness-params`.

    ```bash theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "Qwen3.6-35B-A3B", "base_url": "https://your-judge-endpoint/v1", "api_key": "sk-…"},
        "category": ["science", "geography"]
      }' \
      --harness-params '{
        "tools": ["search", "visit"],
        "max_iterations": 40,
        "serper_api_key": "your-serper-key",
        "jina_api_key": "your-jina-key"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Evaluate the full task set. `--benchmark-params` only needs the judge model `judge_model`; use `--task-concurrency` to raise cross-task concurrency.

    ```bash theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "Qwen3.6-35B-A3B", "base_url": "https://your-judge-endpoint/v1", "api_key": "sk-…"}
      }' \
      --harness-params '{
        "serper_api_key": "your-serper-key",
        "jina_api_key": "your-jina-key"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>
</Tabs>

## Outputs

A run produces two kinds of results, both under `results/browsecomp_zh/<model>/<run>/`: **aggregate metrics** (`summary.md`, overall performance) and **per-task details** (`details/`, per-task grading).

### Aggregate metrics (summary.md)

`summary.md` summarizes the overall performance of the run, in two parts — a run overview and the metrics.

**Run overview**

| Field       | Meaning                                                                                                                                                             |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Model`     | The model-under-test id                                                                                                                                             |
| `Total`     | The total number of loaded tasks                                                                                                                                    |
| `Evaluated` | The number of tasks evaluated (should normally equal `Total`)                                                                                                       |
| `Error`     | The number of tasks that errored during running or judging (`RUN_ERROR`); a value greater than 0 means those tasks produced no valid grading and need investigation |

**Metrics**

There is a single headline metric, **`accuracy`**: the share of tasks judged correct. A task counts as correct (scored 1, otherwise 0) if and only if the judge returns verdict **A**; `accuracy` is the average over all tasks.

### Per-task details (details/)

Each task has one JSON file, in which the judge's grading for the task is recorded under the `extra.scoring` field:

| Field             | Meaning                                                      |
| ----------------- | ------------------------------------------------------------ |
| `evaluation_type` | Fixed as `llm_judge`                                         |
| `correct`         | Whether the task is finally judged correct (judge verdict A) |
| `model_answer`    | The final answer produced by the model under test            |
| `ground_truth`    | The reference answer                                         |

Only the parsed verdict is persisted here; the answer under test and ground truth are kept for tracing, while the full trajectory is written alongside in the same task file.
