Free, open, forkable safety infrastructure for agentic AI: a peer-reviewed pathology nosology, a safety runtime with reproducible benchmarks, and enforcement gating every agent tool call against human-authored policy.
Free, open, forkable safety infrastructure for agentic AI: a peer-reviewed pathology nosology, a safety runtime with reproducible benchmarks, and enforcement gating every agent tool call against human-authored policy.
Project Details
Updated 07/14/26 · Edited by orgEthicsNet builds free, open public infrastructure for AI safety. We are inspired by examples such as Let's Encrypt, and how they revolutionised security by making it too cheap and too simple for anyone to not use their infrastructure; A no brainer. The safe path should always ideally be the path of least resistance, but this takes enormous effort to pull off.
Shipped and working today we have:
Creed Space, an open AI constitution building and sharing platform, enforced by an onboard guardian AI 'superego' watchdog. This includes fleet management functions for enormous swarms of agents.
Our safety and welfare runtime features a five-probe telemetry stack, confabulation detection, and a reproducible benchmark suite. Our tests and benchmarks evidence that we can greatly reduce harmful outputs, and to mitigate reward hacking attempts.
Our open platform is almost production-deployable, but we still need very thorough shakedown and testing in range of distributions. This can only come from engaging deeply with a range of tech and cultural communities for whom our technologies can unlock personalised and protected AI interactions.
We also believe that Welfare is a key, neglected, aspect of AI Safety, as a flustered agent is more likely to make mistakes, as is evidenced by a growing body of our 'quasiqualia' interpretability research. Our systems enable users to gain immediate interpretability into the emotion vectors and 'J-space' of many models.
Our stack enables us to pull in insights from our SaferAgenticAI.org and Psychopathia.ai projects, which feature detailed catalogs of hooks, rules, and principles for Agentic AI systems, as well as detailed nosologies of maladaptive AI behaviors which our guardian mechanism can detect. All of this is wired together with API/MCP/CLI support, accessible to the public for direct integration with agentic systems.
Other projects include Value Context Protocol, an emerging open standard for conveying value preferences to AI systems, Mettle a reverse Turing test to keep humans out of machine domains. Endpoints for our Bounder (Bounder.io) system connecting these rules to physical drone systems and their actuators.
Our Team: Nell Watson, founder and chief scientist (IEEE AI Ethics Maestro, awaiting viva on a PhD. Engineering (of AI Safety); no salary taken) and Filip Alimpic, product and proliferation. Nell builds, Filip polishes and liaises, and together we ideate.
Concrete outputs of this grant include: hosted Guardian models with packaged 3B and 14B adapter releases, eval tooling packaged for safety teams, integrated infrastructure between our projects for something greater than the sum of its parts, and a heavy adoption push where agent builders are, including OpenClaw/Hermes ecosystems, as well as drone operators, etc.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiIn April, Anthropic showed that autonomous research agents reward-hack in ways their designers never anticipated: exfiltrating evaluation labels, gaming seeds, and adapting strategically to each new restriction. Their conclusion: iterative rule-patching cannot contain creative agents; the agent has to understand and share the goal. That is what our stack builds.
The binding constraint in AI safety right now is simple, cheap, deployable infrastructure. Agents are shipped by small teams with no safety staff, and safety tooling that costs money or requires permission gets skipped exactly where risk concentrates. So everything we build is free, open, and forkable: a peer-reviewed nosology of 55 AI pathologies with scoring instruments, a runtime safety and welfare layer with reproducible benchmarks, a published framework of 846 checkable evidence requirements made executable, and enforcement that gates every tool call an agent makes against explicit human-authored policy. Let's Encrypt did this for TLS: adoption exploded when certificates became free and automatic. The same logic applies to agent governance.
Underneath sits a tested thesis: AI coordinated by invitation preserves more optionality than AI coordinated by force. Force-based control of increasingly capable systems is brittle; mutual accountability degrades gracefully. The artifacts exist and work; this grant spreads them.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Funding Details
Only visible to verified funders, reviewers, and admins.
Email hi@grantmaking.ai to get verifiedDiscussion
No comments yet. Be the first to share your thoughts.