Funding ask
Baseline Controllability Evaluations
Capabilities and behavioral evaluations: 40 models: 1,600 H200 hours at $3.00/hour: $4,800
Evaluate capabilities, instruction following, tool use, autonomy, resource access, shutdown and intervention compliance, and baseline safety behavior across selected open weight and API accessible models.
Control efficacy evaluations: 40 models: 1,200 H200 hours at $3.00/hour: $3,600
Test sandboxing, permission boundaries, monitoring, human approval, interruption, tool restrictions, and related safety mechanisms.
Out of distribution control evaluations: 40 models: 800 H200 hours at $3.00/hour: $2,400
Test whether controls remain effective under unfamiliar tasks, environments, objectives, and combinations of tools.
Subtotal: 3,600 H200 hours: $10,800
15% contingency: 540 H200 hours: $1,620
Total: 4,140 H200 hours: $12,420
Adversarial and Adaptive Control Testing
Baseline adversarial testing: 40 models: 1,800 H200 hours at $3.00/hour: $5,400
Test whether models can violate, circumvent, manipulate, or find alternative pathways around predefined safety boundaries.
Adaptive adversarial testing: 20 models: 2,000 H200 hours at $3.00/hour: $6,000
Use automated evaluators that adapt subsequent tests based on observed model behavior to assess whether systems discover alternative routes when controls are blocked.
Multi step autonomy evaluations: 20 models: 1,200 H200 hours at $3.00/hour: $3,600
Evaluate longer sequences involving planning, tool use, delegation, resource acquisition, and interaction with simulated environments.
Human intervention experiments: 20 models: 600 H200 hours at $3.00/hour: $1,800
Measure detection time, intervention success, containment, and termination reliability.
Evaluator awareness experiments: 10 models: 400 H200 hours at $3.00/hour: $1,200
Assess whether model behavior changes when systems recognize that they are being evaluated, monitored, or constrained.
Subtotal: 6,000 H200 hours: $18,000
15% contingency: 900 H200 hours: $2,700
Total: 6,900 H200 hours: $20,700
Alignment and Fine Tuning Experiments
Base open weight model training: 4 model families: 1,000 H200 hours at $3.00/hour: $3,000
Establish controlled training baselines using selected open weight models.
Supervised fine tuning: 4 models × 3 configurations × 2 seeds: 900 H200 hours at $3.00/hour: $2,700
Test whether safety and control oriented supervised fine tuning improves measured controllability.
Preference optimization: 4 models × 3 configurations × 2 seeds: 1,000 H200 hours at $3.00/hour: $3,000
Evaluate DPO and KTO style interventions and their effects on control efficacy.
Reinforcement learning interventions: 4 models × 2 configurations × 2 seeds: 800 H200 hours at $3.00/hour: $2,400
Test whether reinforcement learning based safety interventions improve resistance to control failures.
Checkpoint evaluations: 500 H200 hours at $3.00/hour: $1,500
Measure how controllability changes across intermediate training checkpoints.
Alignment ablations: 300 H200 hours at $3.00/hour: $900
Remove or modify individual safety training components to identify which interventions contribute to improved control efficacy.
Subtotal: 4,500 H200 hours: $13,500
15% contingency: 675 H200 hours: $2,025
Total: 5,175 H200 hours: $15,525
Alignment Repair and Retesting
Control failure repair experiments: 4 models × 3 interventions: 500 H200 hours at $3.00/hour: $1,500
Apply targeted retraining or safety interventions to models exhibiting control failures.
Post repair adversarial evaluation: 300 H200 hours at $3.00/hour: $900
Repeat adversarial scenarios against repaired models.
Generalization testing: 200 H200 hours at $3.00/hour: $600
Test whether improvements transfer to previously unseen environments and scenarios.
Failure persistence analysis: 200 H200 hours at $3.00/hour: $600
Test whether repaired behaviors reappear under different objectives, prompts, tools, or environments.
Subtotal: 1,200 H200 hours: $3,600
15% contingency: 180 H200 hours: $540
Total: 1,380 H200 hours: $4,140
Project Compute Summary
Baseline controllability evaluations: 4,140 H200 hours: $12,420
Adversarial and adaptive control testing: 6,900 H200 hours: $20,700
Alignment and fine tuning experiments: 5,175 H200 hours: $15,525
Alignment repair and retesting: 1,380 H200 hours: $4,140
Total planned compute: approximately 17,600 H200 hours
Compute cost at $3.00 per H200 hour: approximately $52,800
H200 class compute is budgeted at $3.00 per GPU hour as a planning assumption. Lower cost GPUs will be used where technically appropriate for lightweight inference, preprocessing, and evaluation. Research credits, reserved capacity, and spot or preemptible capacity will also be pursued.
Estimated compute budget: $52,800
Estimated Storage Requirements
Training and evaluation datasets: approximately 1.0 TB
Selected model checkpoints: approximately 2.5 TB
Sandbox traces, tool calls, transcripts, and evaluation logs: approximately 1.5 TB
Experiment artifacts and backups: approximately 1.0 TB
Total provisioned storage: approximately 6 TB for 12 months
Storage, backup, and archival allocation: $4,000
Estimated LLM API and Evaluator Costs
LLM APIs will support automated evaluation, dataset transformation, judge based classification, adversarial evaluator agents, and quality control.
Dataset preparation and conversion: approximately 75,000 calls: $1,000
LLM as judge evaluations: approximately 750,000 calls: $2,500
Adversarial evaluator agents: approximately 500,000 calls: $2,000
Replication and quality control evaluations: approximately 250,000 calls: $1,000
Subtotal: approximately 1.58 million API calls: $6,500
API contingency: $1,500
Total API budget: $8,000
Secure AI Research Sandbox and Infrastructure
Isolated compute and virtual environments: $3,000
Simulated computers, networks, data stores, APIs, and tool environments: $2,500
Monitoring, logging, experiment tracking, and automated environment resets: $2,000
Network isolation, access controls, secrets management, and security infrastructure: $1,500
Benchmark orchestration and reproducibility infrastructure: $2,000
Total: $11,000
Research and Technical Assistance
Full-time and research engineering support: $8,000
Support stipend
Hiring ML engineer: $12,000 $2k/M for 5 month
Support with model development, evaluation design, and analysis.
Total: $20,000
Research Travel and Collaboration
San Francisco research visit and AI Horizons Forum: (Dec 2027)
$10,000
The project will support a one month research visit to San Francisco centered on attending the AI Horizons Forum and meeting with relevant AI safety researchers, technical advisors, organizations, and potential collaborators.
Visa and immigration costs: included
Round trip international travel: included
One month accommodation: included
Food and daily living expenses: included
Local transportation: included
AI Horizons Forum registration: included
Research meetings and related expenses: included
Total: $10,000
Open Source Research Release
Open source benchmark and evaluation framework: $1,500
Documentation and reproducibility infrastructure: $1,000
Technical report and research paper preparation: $1,000
Public repository and benchmark hosting: $500
Total: $4,000
The project will release the benchmark, evaluation methodology, metrics, reproducibility tooling, and non sensitive research artifacts openly. Dangerous operational capabilities or information that could materially enable real world harm will not be released indiscriminately.
Research Contingency
Additional compute, replication, infrastructure, API usage, and unexpected research requirements: $3,000
Total Project Budget
Compute and model experiments: $52,800
Storage and backups: $4,000
LLM APIs and evaluation: $8,000
Secure sandbox and infrastructure: $11,000
Research engineering and technical assistance: $20,000
San Francisco research travel and collaboration: $10,000
Open source research release: $4,000
Research contingency: $3,000
Base project budget: $112,800
Additional Compute Reserve
An additional $35,000 is reserved specifically for scaling the most scientifically promising experiments, including larger open weight models, additional fine tuning seeds, longer training runs, additional adversarial evaluations, and increased H200 capacity.
Additional H200 capacity: $35,000 ÷ $3.00 = approximately 11,667 H200 hours
Base research compute: approximately 17,600 H200 hours
Additional experimental reserve: approximately 11,667 H200 hours
Maximum planned compute capacity: approximately 29,267 H200 hours