Core problem: AI safety evaluations are immature (see Apollo's We need a science of evals), yet they inform high-stakes decisions like Responsible Scaling Policies. Extensive relevant expertise exists in social sciences and STEM fields that have long studied the behaviours evals seek to measure (deception, bias, power-seeking, CBRN threats), but technical barriers keep those experts out: to contribute, they must either become AI safety engineers or partner with one. Partnerships are often possible through highly competitive programmes, which also favour technical AI safety engineers for empirical research (e.g. the use of standardized coding coding tests that most social scientists are not trained to complete)
Accessible tooling removes these barriers and scales the efforts of existing engineers.
Project goal: accelerate maturation of AI safety evaluations by democratising access and contribution to AI safety evals through accessible tooling for AI safety, such as no-code evals (including red teaming) tooling.
Compressed: accessible tooling → broader expert participation and greater throughput → more mature evaluations → better safety decisions → reduced AGI/TAI risk.
This rests on assumptions we have investigated already: that technical barriers (not lack of interest) are the primary bottleneck; that domain expertise transfers meaningfully to evaluating analogous behaviours in AI; and that broader participation raises quality rather than diluting it.
We built a quick prototype with limited features that we temporarily made available, and we hosted a hackathon to test our assumptions. The feedback from the hackathon has been positive and signals the usefulness of the tool in contributing to reducing x-risks through evals and red teaming by domain experts. Example feedback we received: