Collaboration of researchers running open-world, real-world evaluations of frontier AI capabilities.
Collaboration of researchers running open-world, real-world evaluations of frontier AI capabilities.
People
Updated 06/29/26 · By grantmaking.aiPrinceton University, Cornflower Labs
Princeton University
Princeton University
Princeton University
Org Details
Updated 06/29/26 · By grantmaking.aiGenerated by AICRUX (Collaborative Research for Updating AI eXpectations) is a collaborative effort to regularly run “open-world evaluations” of frontier AI capabilities—evaluations built around long-horizon, messy, real-world tasks where success cannot be cleanly specified or automatically graded. The project positions these evaluations as complementary to conventional benchmarks, aiming to capture agent behaviors and constraints that are hard to standardize or reproduce, and to provide richer evidence about what agents can and cannot do in realistic conditions. CRUX’s stated evaluation format includes a real-world task, an agent scaffold that could plausibly enable completion, detailed log analysis, and a public write-up that incorporates interpretations from collaborators with diverse perspectives. The project also plans a regular cadence of releases, aiming to publish new evaluations every 1–2 months, with upcoming work spanning domains including AI R&D tasks.
CRUX #1 tested whether an AI agent could autonomously develop and publish an iOS application. In this experiment, the team gave the agent access to accounts and infrastructure needed to interact with real external systems (including an Apple Developer account and a Mac virtual machine) and set a concrete success criterion: getting an app published on the App Store. The agent was responsible for nearly all steps involved in shipping the app—coding, building, preparing metadata, drafting and hosting a privacy policy, submitting for review, and responding to feedback—while humans handled steps required by platform policy (e.g., account setup and final release approval). The experiment succeeded: the agent built and published an iOS app with minimal human involvement and an overall cost of about $1,000, with much of the expense driven by extended monitoring during the App Store review wait rather than by development itself.
Beyond the headline success, CRUX #1 documents practical limitations and failure modes that emerge in real workflows. The write-up notes issues such as fabricated or missing information (e.g., a fictional phone number for review contact), credential management problems that triggered a manual intervention, and quality gaps in UI-centric deliverables (like a listing screenshot with formatting errors and an app feature that did not work as intended). CRUX also emphasizes the relevance of external “platform frictions” (e.g., review timelines and submission rules) and how scaffolding choices can have large cost implications (for example, frequent heartbeat-driven status checks). To support transparency and cumulative learning, CRUX releases the full logs from the evaluation and explicitly invites the broader community to explore them and share findings. The team also describes a responsible-disclosure step—contacting Apple’s product security team ahead of publication—motivated by the possibility that near-future agents could scale similar app-submission behavior dramatically.
Theory of Impact
Updated 06/29/26 · By grantmaking.aiGenerated by AICRUX’s theory of change is that open-world evaluations—long-horizon tasks in real-world environments, analyzed qualitatively through detailed logs—provide more realistic and decision-relevant evidence about frontier AI agent capabilities than benchmarks alone. By eliciting “upper-bound capabilities” in messy settings and documenting where agents succeed or fail (including errors, unintended actions, and operational frictions), CRUX aims to generate early warning signals about capabilities that may soon become widespread. Publishing write-ups and releasing full logs is intended to enable broader community scrutiny and cumulative learning, improving collective understanding of risks and limitations. In cases where evaluations reveal security-relevant scaling risks (e.g., the prospect of agents submitting many apps), CRUX’s approach also includes responsible disclosure to affected stakeholders, supporting earlier mitigation by platforms and other actors.
Projects
Updated 06/29/26 · By grantmaking.aiDiscussion
No comments yet. Be the first to share your thoughts.