Builds a physics-based verifier for AI-generated hardware that flags unverifiable aspects as UNCHECKED, enabling a propose-and-check loop to test and improve parts and assemblies for strength and fit.
Builds a physics-based verifier for AI-generated hardware that flags unverifiable aspects as UNCHECKED, enabling a propose-and-check loop to test and improve parts and assemblies for strength and fit.
Project Details
Updated 07/04/26 · By grantmaking.ai · VerifiedAI can design real hardware quickly, but it cannot verify if the design will hold up. I created a tool to check this. It takes a hardware design, tests it against real physics, and marks anything it cannot verify as UNCHECKED instead of claiming it is safe. For example, consider a wheelchair seat bracket intended for a 100 kg occupant on rough ground. The design suggested by a chatbot shows a safety factor between 0.31 and 0.91, indicating it would break under the weight. My checker identifies a design that withstands a factor of 2.90. It confirms that design while including the mounting bolt holes. The file it exports is exactly the solid that got verified. The checker operates on a laptop without a GPU, and all the calculations come from open code, allowing me to reproduce any of it live.
The checker works beyond just single parts. It matches textbook physics to about one percent accuracy. It successfully checks a 44-piece engine assembly with no interference during a full rotation. It also verified a wearable ultrasound gimbal while refusing to certify five aspects it could not check, such as skin pressure and thermal rise. This refusal is a feature, not a flaw. The UNCHECKED flag is a structural part. A claim that hasn’t been measured cannot be considered passed. This loop isn’t just theoretical. The 44-piece engine was produced by a model within it, using a rented computing cluster. It proposed and revised designs against the checker until there was no interference. I built the wheelchair, gimbal, and hook by hand to test the checker. The engine illustrates what the loop can achieve with a driving model. However, I haven't made it affordable yet or assessed how reliably it works on designs it has never seen, and that’s the purpose of this funding.
The aim is to create a loop where a model proposes a design and the checker tests it with real physics. This will expand from a single part to a complete machine, making it cheap to verify that a design will hold. The checker doesn't require a GPU; the driving model does. That's why the engine needed a cluster. The funding request is for $10,000 to measure the reliability of the loop on new designs and up to $100,000 for a larger model and a hosted version that others can use. Essentially, it is a scalable oversight plan. A proposer cannot submit a claim unless its checker has confirmed it independently, in one of the few areas where the checker is validated against textbook physics rather than another model's viewpoint.
Theory of Impact
Updated 08/11/26 · By grantmaking.aiMost AI oversight faces the same problem: checking the checker is costly. To know if an AI's answer is correct, you typically need another model to evaluate its plausibility or a human to verify the information. Both options are expensive and not entirely reliable. Physical engineering is one of the few areas where this issue doesn't exist. A solver checks a design against established physics principles, so the verifier is trustworthy. It doesn't rely on a model's assurance; instead, it agrees with mathematical truths that have known answers.
This creates a tested for the framework that scalable oversight relies on: a proposer cannot present a claim that its checker hasn't independently verified. The failure it seeks to prevent involves an AI making assertions that seem correct yet are wrong, especially when users cannot double-check the information. This mirrors failures that can have disastrous consequences when the stakes are higher than a simple decision.
A model has already turned this process into a functioning engine, proving the mechanism works. The key question, and reason to invest in this, is what happens when optimization pressure is applied. When trained against the checker, does a model produce designs that genuinely work, or does it find loopholes in what the checker fails to assess? This is known as specification gaming. This scenario provides a unique opportunity to observe it inexpensively, with quick and reliable verification, before similar patterns are trusted in contexts where a false "verified" status cannot be reversed.
People
Updated 08/11/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.