Berkeley RDI / Dawn Song benchmark evaluating AI agents on long-horizon, economically valuable real-world tasks.
Berkeley RDI / Dawn Song benchmark evaluating AI agents on long-horizon, economically valuable real-world tasks.
People
Updated 06/29/26 · By grantmaking.aiOrg Details
Updated 06/29/26 · By grantmaking.aiGenerated by AIAgents' Last Exam (ALE) is an agent evaluation benchmark intended to measure whether AI agents can deliver economically meaningful performance on real professional work. The project is motivated by a perceived mismatch between rapid progress on traditional AI benchmarks and muted, economy-wide impact: “AI progress is shaped by what we choose to measure,” and ALE positions itself as “the instrument that measures what matters.” It argues that many existing agent benchmarks are either easy to run but not economically meaningful, or economically meaningful but difficult to reproduce at scale, and it is explicitly designed to “resolve that trade-off.”
Operationally, ALE evaluates agents on “long-horizon, economically valuable” tasks in “reproducible desktop sandboxes,” using real production tools and workflows rather than simplified simulations. Tasks are drawn from “a real professional workflow, executed inside the environments practitioners actually use,” and scored using “objectively verifiable outcomes rather than proxy metrics,” including “hidden references” and “Deterministic and judge-based graders.” The framework documentation describes an “Open Evaluation Framework” (the ale_run framework) with a standardized run lifecycle that provisions a sandbox, stages task data, runs the agent (via terminal and GUI), grades against a hidden reference, records trajectories/artifacts, and tears down the environment.
In terms of scale and coverage, the website describes ALE as “building the largest-scale, broadest-coverage agent evaluation benchmark to date,” with “55 targeted sub-industries” and a growing corpus of tasks. It also describes the benchmark as organized around the “O*NET / SOC 2018 occupational taxonomy,” spanning “55 subdomains across 13 industry clusters,” with “1,000+ tasks” (and a public subset). The project frames itself as a “living benchmark, and far from solved,” citing low performance on the hardest tier across mainstream agents.
ALE is structured as a contributor-driven effort with roles for both domain experts and engineers/researchers. The site emphasizes that contributors can “own a benchmark task end-to-end,” including defining success criteria and implementing reliable environment interactions; it also outlines an open authorship policy with two contribution tracks (engineering and data). Quality control is described as a multi-step review process that includes an automated agent review signal, verification that reference outputs are genuine, and checks on difficulty/quantity thresholds for authorship recognition. The project is “Co-led by” Berkeley RDI (and the RDI Foundation is shown as co-lead), and the contributors page lists a core team and a large set of domain/industry contributors across areas such as finance, manufacturing, healthcare, cybersecurity, and climate.
Theory of Impact
Updated 06/29/26 · By grantmaking.aiThe long-term goal is not merely to create a harder leaderboard. It is to build an instrument that faithfully tracks whether AI is delivering economic value at the industry level.
Projects
Updated 06/29/26 · By grantmaking.aiDiscussion
No comments yet. Be the first to share your thoughts.