grantmaking.ai Launch Round
I am working on Halo-2.0, a persistent autonomous AI agent. Give it a goal and walk away. It plans the work, uses a browser and other tools, reads their output, writes/runs code when necessary, adjusts itself in case a certain step fails and continues moving toward the goal without approval after each step.
Working demonstration: https://youtu.be/lGcZGRFRXQ8
Project website: https://www.oscerra.space
Halo-2.0 is currently in closed alpha. The whole system already works end-to-end, but at some point I've realized that testing it on my own tasks is not fully sufficient. This grant application is not about releasing a product publicly or scaling it. This grant application is a 12-week external testing program to study how it behaves on tasks chosen by people other than its developer.
The central design question of my work is quite straightforward: What should an agent do when it believes a step succeeded, but the result shows that it did not?
Many autonomous agent systems incorporate planning, action, and judgment into one model loop. Model says that the file is edited, the page is understood, the build is done and/or the answer was found and then agent goes further. My approach is aimed to challenge such assumption. In case of executing an action, system checks what has really changed in the browser, file system, the output of commands or in results of using a tool. In case of a mismatch between the expected change and reality – step is considered incomplete, agent reconsiders the situation, modifies its approach, asks for help and even stops if necessary.
The purpose of the grant will be to:
-
Isolate the Halo agent into a testable version with all the precautions taken (
isolated sessions, bounded permissions etc.)
-
For $5,000, a smaller cohort group of 8-10 testers will be supported for 12 weeks, and around 50-60 traces of tasks will be completed. For $10,000, a larger group of 15-20 testers will be supported, along with 100 or more task traces, some testing
-
Capture and replay at least 100 completed task runs including unsuccessful attempts, recovery from them, repeated mistakes and early stops.
-
Transform those runs into the taxonomy of long-horizon autonomous agent errors and small evaluation dataset to test possible solutions to the errors.
-
Publish a report of what errors occurred, what interventions were helpful, which were not and what are the safe limitations of Halo.
Codebase remains proprietary. Public deliverables will include the evaluation procedure, anonymized failure cases and findings for others interested in developing or evaluating autonomous systems.
I am Mashrikain Mazdi; a 16-year-old self-taught programmer from Barisal, Bangladesh. I began to develop web automation systems at the age of 13, realizing that the world described in the documents was not the world encountered by software on the web. I've developed Halo-2.0 completely solo, including those of planning, web interaction, model routing, memory, programming, code reviewing, and recovery. External testers will participate in both versions of the project (minimum funding and ideal funding). The ideal version would also include limited review from independent technical reviewers.
Minimum version — $5,000
- $2,500 — Model and API usage: multi-provider inference, evaluation calls, retries, and targeted stress-testing for a reduced external cohort.
- $750 — Cloud compute: isolated browser and coding environments for 8–10 technical testers.
- $500 — Logging, storage, and replay: task traces, screenshots, tool outputs, and reproducible failure cases.
- $500 — Tester incentives: modest compensation for sustained, high-quality participation.
- $500 — Security, onboarding, and evaluation tooling: bounded permissions, session isolation, basic onboarding, and an initial failure-labeling system.
- $250 — Contingency: provider price changes, unusually long runs, and unplanned infrastructure costs.
At the minimum amount, I would run a meaningful reduced 12-week cohort with approximately 8–10 technical testers and 50–60 complete task traces. The outputs would include an initial structured failure taxonomy, anonymized failure cases, and a pilot report. Independent technical review and the full evaluation-set release would be reserved for the ideal version.
Ideal version — $10,000
- $4,000 — Model and API usage: longer tasks, repeated trials, multiple model routes, retries, and evaluation calls.
- $1,500 — Cloud compute: isolated browser and coding environments for 15–20 testers, with greater concurrency.
- $1,000 — Logging, storage, and replay: a larger durable corpus of task traces, screenshots, tool outputs, and reproducible failure cases.
- $1,500 — Tester incentives and external review: sustained participation from serious technical testers, plus limited independent review.
- $1,500 — Security, onboarding, and evaluation tooling: stronger permission boundaries, safer test environments, structured failure labeling, and a cleaner testing interface.
- $500 — Contingency: provider price changes, unusually long runs, and unplanned infrastructure costs.
The ideal amount would support the full 12-week cohort of 15–20 testers, at least 100 complete task traces, a formalized failure taxonomy, a small evaluation set, and a public report. No part of either budget is for marketing, paid acquisition, hardware, or general commercial runway.
Detailed budget and 12-week work plan: Halo 2.0 | Detailed Budget - Google Sheets