grantmaking.ai Launch Round
Min :
- 1 FTE for the development of the solution for 6 months + tools (Claude code)
Ideal:
18 months runway for
- 2-3 zkp engineer
- 2 GPU kernel engineer
- 2 distributed training engineer
- 1 software eng.
- 1 CTO
- 1 CEO
- 1-2 Ops
Development of a sampling based zero knowledge proof for frontier AI training with 10% overhead.
Endorsements made here support General Purpose AI Policy Lab.
Development of a sampling based zero knowledge proof for frontier AI training with 10% overhead.
Endorsements made here support General Purpose AI Policy Lab.
The current problem:
No verification mechanism exists so far to provide proof of training, i.e. specific properties of a frontier AI training run such as its total FLOPs budgets or specific properties regarding its dataset or training procedure.
Such mechanisms are critical components for international coordination over AI development and enforcing national regulations.
While solutions do exist for proving inference using cryptographic solutions known as zero-knowledge proofs, such solutions were considered intractable to prove training so far as they are estimated to generate a 10^5-10^6 computation overhead.
Our approach:
Our approach (developed at the GPAI Policy Lab)using a different paradigm aims to produce local proof of training by via sampling with less than 10% overhead.
What we do differently:
We leverage network observations (via hardware network TAP or retrofitted smartNIC/DPU) and zero-knowledge virtual machines to produce local proofs instead of monolithic proof of training.
This avoids the constraint explosion that monolithic proofs produce by trying to prove the entire training within one unique proof.
Built upon an existing demo, our first funding round targets the development of three prototypes over 10 to 18 months (see table 2 below for detailed timelines).
Those prototype specs are designed to de-risk the technical roadmap as soon as possible by “failing as fast as possible”: a successful prototype 1 allows to parallelize the development of prototype 2a and 2b, and successful prototypes 2a and 2b allow the development of the future MVP.
A failure anywhere in the development chain would help identify critical rework early during the development cycle.
Deliverables
Demo: proof of forward / backward over 1 layer of a toy model without non-linearity (Linear Attention → MLP, no activation)
Prototype 1: toy complete transformer
Proof of forward or backward over 1+ layers of a complete toy model at toy dims (e.g. d_model=256, n_heads=8, seq=128, 2 layers).
Prototype 2A: real-shape transformer block
Proof of forward and backward over a single transformer block at frontier dims (d_model ≈ 2048, n_heads ≈ 16, seq ≈ 1k to 2k)
Prototype 2B: distributed toy with simulated network observation (no hardware)
Proof of distributed training across 4-8 H100 handling distributed communications within the zkVM computation graph (API-level capture, not wire-level)
[Optional] PyTorch → arch_spec converter
Automatic conversion of PyTorch scripts into the arch_spec language of the zkVM. Without it, arch_spec must be handcrafted to match the training script.
Phase 3
MVP : not covered by this grant
Our objective is to develop three prototypes within a 18 months timeline that could be merged into a MVP.
The MVP should be deployable in ~12 months allowing for enforcement of internal agreements in less than 36 months.
International agreements over the development of frontier AI can be grounded in trustworthy verification mechanisms for training, significantly improving their chances of success (international agreement between rivals without verification mechanisms are very hard to settle and enforce).
Team Member
CEO
Min :
Ideal:
18 months runway for
No comments yet. Be the first to share your thoughts.