Skip to main content
Choose a Model, Benchmark, Harness, and Environment, validate one task, then tune the run. If you have not completed an evaluation yet, start with the Quick Start.

Choose the Four Components

The command’s positional order is Benchmark, Harness, Model:
For example, with the Model variables and Docker setup from the Quick Start, run one SWE-bench Verified task:
sample_ids selects the task, and concurrency 1 makes the initial run easier to inspect. Once it completes, use the command builder to configure a larger evaluation.

Choose an Interface

Both interfaces use the same runtime and run configuration. Save repeated settings in a YAML file and load it with --config; see configuration files and precedence. The model under test is supplied through the command, an orchestration request, or the SDK.

Adjust the Run for Your Goal

Use the component pages above for model connectivity, task selection, agent behavior, and sandbox settings. Use the shared guides below for controls that apply across components: