Runtime Status
When to Use
Use ScreenSpot when you need to measure GUI grounding behavior with the task assumptions described by this benchmark. For large or remote benchmarks, prefer benchmark recipes so images, workspaces, and provider-specific defaults come from task metadata instead of manual CLI flags.Parameters
Common parameters for this benchmark include:categorysample_idsagent_typemax_concurrency
sample_ids follow Benchmark Parameters. Pass the ScreenSpot-specific category field in the same --benchmark-params object. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.
Run Example
Outputs
Per-task details are written to the run directory’sdetails/ subdirectory. The same run directory contains the aggregate views summary.md and metrics.json.
Notes
Setcategory to desktop, mobile, web, or all.