grantmaking.ai Launch Round
We believe many failure modes of contemporary AI systems (LLM, VLMs, etc.) are difficult for people to anticipate precisely because the systems are trained and deployed under non-human conditions. Internet-scale next-token training with effectively superhuman memory for textual regularities and weak grounding in action can produce competent behavior, but it does not reproduce the developmental constraints that shaped human cognition. Human intelligence is shaped by the constraints of limited memory, limited sensing and actuation bandwidth, limited processing speed, partial observability, specific spatiotemporal scale, and the need to rely on tools and other people for increasing agency and cognitive offloading.
We propose to adapt the learning-progress view of curiosity as a single general objective under realistic environmental constraints to imbue frontier systems with human-like ways of solving tasks and efficiency. Under this paradigm intrinsically valuable patterns are neither random nor trivial, but those for which the agent can currently improve compression or prediction. As a proof-of-concept study, we want to test this approach in one of the hardest AGI challenges: ARC-AGI-3. We believe the key to closing the gap between human and AI performance lies in endowing frontier models with curiosity-as-prediction-progress and awareness of constraints: formalized notions of the observation and action interfaces, time, and energy.
Our core framework encompasses 3 distinct models: the actor, the world model, and the goal model — these loosely map to different functions of the human brain. Each of these is instantiated as a small VLM. We use the term “agent” as the ensemble of these 3 components. The world model is trained to accurately predict future observations, while the goal model is trained to predict the next optimal action/observation given an implicit understanding of the current goal. The actor (e.g. a VLM) itself is the controller, and it makes use of the world and goal model predictions. The actor decides how much time to spend thinking, whether it should collect simulated observations from the world/goal models, or produce an action. It learns to choose between these efficiently through our objective which incorporates intrinsic motivation and resource constraints (i.e. a notion of time/compute).
Success means agents that are able to finish 100% of the ARC-AGI-3 games with close to human action efficiency. This is a hard target, but we wish to make as much progress as possible in 3 months, open-source all code and results, and write a strong publication on our work and findings.
Participants
Richard Csaky - project lead and research
https://www.linkedin.com/in/richard-csaky/
https://richardcsaky.notion.site/main
Richard’s record spans multimodal foundation models, neural time series, and real time ML systems, with publications in NLP and neuroscience plus recent foundation model scale NeuroAI work. He obtained his PhD from Oxford studying the intersection of human and artificial intelligence. His most recent project was funded by a Foresight Institute AI Safety grant, training a long context generative model on 500+ hours of brain data.
Come Chevalier - engineering
https://www.linkedin.com/in/côme-chevalier-bb5526147/
Holding a Master's degree in engineering, Côme has built strong experience in the automotive industry (Bosch and aiMotive), spanning development and system-level responsibilities in autonomous driving. More recently, he has worked across a range of applied Al topics, including NeRF and Gaussian Splatting as well as web solutions integrating agentic Al components.
In the first part of the competition, we explored various agentic approaches and at the time of the June 30 milestone cutoff we were in the top 10 on the public leaderboard out of more than 1500 participants, our team name: Face of AGI. Our final solution at this date and all our research can be found here: https://github.com/cocomette/artificial-agency.
- Compensation for 2 people full-time for 3 months: $36,000.
- Cloud compute budget: $3,000.
- Used for local model inference, baseline runs, and lora-finetuning.
- Local development hardware: $12,000.
- Two laptops or workstations for full-time development, local inference, storage, and evaluation work.
- Software tooling: $1,000.