Required Content
Document:- Purpose and official source links.
- Supported dataset and evaluator versions with pinned revisions.
- Task count, categories or splits, task browser, license and access requirements.
- Prerequisites, optional dependencies, task images and credentials.
- Recommended official Harness and other compatible Harnesses.
- Supported Environments and provider-specific behavior inferred by Recipes.
- Benchmark-specific parameters, defaults, valid values and selection guidance.
- Benchmark-specific output and metric semantics.
- One real smoke command and one complete evaluation command.
- Known compatibility constraints and official-alignment notes.
Keep the Page Benchmark-Specific
- Link generic
k,avgk,sample_ids, and aggregation controls to the shared Benchmark parameter page. - Link Harness installation, step limits, cost tracking, command timeouts, and Model settings to Harness pages.
- Do not describe the positional Model ID as a Benchmark parameter.
- Do not repeat generic
pass@koravg@koutputs unless the Benchmark defines different semantics. - Omit command parameters whose defaults already produce the intended run.
- Label the upstream alignment path Recommended Harness and alternatives Other optional Harnesses.
- Give every alternative Harness a complete evaluation command, not a command fragment.
Preview and Validate
Add the page todocs/docs.json, keep locale paths symmetric unless maintainers approve staged localization, preview it,
and validate it:
