Skip to main content
ScreenSpot evaluates GUI grounding by asking a VLM agent to identify target regions in screenshots.

Runtime Status

When to Use

Use ScreenSpot when you need to measure GUI grounding behavior with the task assumptions described by this benchmark. For large or remote benchmarks, prefer benchmark recipes so images, workspaces, and provider-specific defaults come from task metadata instead of manual CLI flags.

Parameters

Common parameters for this benchmark include:
  • category
  • sample_ids
  • agent_type
  • max_concurrency
Shared benchmark fields such as k, avgk, and sample_ids follow the conventions in Benchmark Parameters. Pass the ScreenSpot-specific category field in the same --benchmark-params object.

Run Example

Adjust the harness and environment to the supported combination for your branch and deployment.

Outputs

Per-task details are written to results/screenspot/<model>/<run>/details/. Aggregate metrics are written to summary.md in the same run directory.

Notes

Set category to desktop, mobile, web, or all.