Geospatial agent benchmark · none

OpenEarth-Bench / OpenEarthAgent

Also known as: OpenEarth-Bench, OpenEarthAgent, OpenEarth Agent

Evaluates: models with tool access

Unified tool registry with structured tool calls and GIS, spectral, and GeoTIFF tool trajectories.

At a glance

pending verificationTasks
UnclearAgents run real code
NoData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortWhether agents run real code here is not confirmed from public sources, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Independent verification records or projection readiness. It is a training or evaluation framework, not a production verification contract.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Unified tool registry with structured tool calls and GIS, spectral, and GeoTIFF tool trajectories.

A good score does not show

Independent verification records or projection readiness. It is a training or evaluation framework, not a production verification contract.

Arena status

none — Not downloaded. Identified in Axis geospatial verification gap research as a unified tool registry with structured tool calls covering GIS, spectral, and GeoTIFF tool trajectories. Deeper fields pending verification pending source review.

All recorded facts (22 fields)
Evaluates
model+tools
Task families
unified tool registry,structured tool calls,GIS, spectral, GeoTIFF tool trajectories
Task count
pending verification
Difficulty
pending verification
Geography
pending verification
Modality
mixed
Input types
GIS data,spectral data,GeoTIFF rasters,structured tool calls
Output types
tool trajectories
Execution environment
unified tool registry with structured tool calls (pending verification)
Ground truth method
pending verification
Verifier method
structured tool-call and tool-trajectory evaluation
Deterministic verification
yes
Scoring dimensions
tool-call correctness,tool-trajectory evaluation
Aggregation formula
pending verification
Trials policy
pending verification
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
pending verification
Reproducibility status
pending verification (described as a training or evaluation framework, not a product)
First published
2026
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#openearth-bench-openearthagent. See the registry endpoint.