launch shares one scheduler across its requests.
Scale Concurrency Safely
Start with a few representative tasks at concurrency1, then increase to 2 or 4. Watch model latency, error rates, Environment startup time, and memory use. Return to the last stable value if errors increase.
The lower of task concurrency and the provider limit caps actual concurrency. Model quotas and host resources can reduce it further. The creation rate controls startup pacing, not the number of active tasks. For CPU and memory per sandbox, use Environment resource settings.
CLI Syntax
Add--task-concurrency 4 to your existing run command. For a multi-request launch using Docker and Modal, you can set separate provider limits:
Configuration File Syntax
Save repeated settings in a configuration file and load it with--config:
launch orchestration file, shared task_concurrency belongs at the top level; provider mappings remain under runtime. See orchestration fields.
If you configure multiple attempts per task, they share the same concurrency pool. Same-task attempts overlap only when the Benchmark and Harness support isolated execution. See repeated attempts.
Retry Only Transient Failures
Retries are disabled by default. Add--max-retries 2 to allow up to two retries after the initial execution of each logical attempt. To retry selected agent errors, also supply patterns, for example:
An omitted,
null, or empty pattern list retries only FATAL issues. Patterns match individual issue messages or codes, not tracebacks. A failure caused by an incorrect endpoint, missing credential, or invalid configuration needs a configuration fix; repeating the same request will not resolve it.
A retry replaces work within the current logical attempt; it does not add another sample to the score. Completed sibling attempts remain saved. Evaluation-only retry is possible when complete scoring inputs are available in none or fresh mode; reuse mode reruns the attempt. See retry details for the recorded history and score validity for unresolved failures.
Output and Reuse
Name a New Run
Choose an experiment group and run ID by adding--run-name ablation --run-id baseline to your run command. With the default result root, the output path is:
Model, Benchmark, and Harness IDs are normalized and joined into one directory name. In
launch, each request’s name replaces that combined directory; see request naming.
Open a completed run with the result viewer:
Resume an Interrupted Run
Add--reuse to the same evaluation command to continue from the latest compatible run, or --reuse 20260806_120000 to select a source run ID. Keep the same result root, run-name group, and Model/Benchmark/Harness selection so AgentCompass can find the source.
Reuse writes a new run and preserves the source directory. It checks the Benchmark, task IDs, attempt plan, and saved data before reusing completed results or recovering pending work. For launch, reuse searches within each request’s named output directory.
The current retry budget applies to unresolved failures; previous retry counts remain history. You can adjust evaluation timeouts or resources when recovering pending evaluation, but reuse does not restore a live agent process or sandbox. See reuse validation and evaluation checkpoints for compatibility requirements.
To rerun pending work instead of restoring its evaluation inputs, add
--no-checkpoint-resume:
true. It does not force already completed results to run again. For launch, set it under defaults.runtime or requests[].runtime; in the SDK, pass checkpoint_resume=False.
Keep Environments for Debugging
Add--keep-environment to your run command when you need to inspect task or evaluator sandboxes after a failure. AgentCompass closes Harness sessions but skips Environment cleanup.
Retries and multiple tasks can leave several resources active. Release them afterward with the provider’s tools. If you only need local copies of files, save artifacts instead.
Logs and Progress
For CI or redirected output, add--progress plain. To see more console detail, add --log-level DEBUG. Console and file log levels are independent:
auto shows a live progress view in an interactive terminal and text progress otherwise. none disables terminal progress; AgentCompass still saves progress, logs, and task results. Follow Diagnose a Failed Run to inspect them.