grantmaking.ai Launch Round
Redarc Labs Lab Motivation: We are an AI safety Lab focused on interpretability under adversarial conditions. Red in our name stands for all the misalignments, jailbreaking and other dangerous…
Redarc Labs
Lab Motivation:
We are an AI safety Lab focused on interpretability under adversarial conditions. Red in our name stands for all the misalignments, jailbreaking and other dangerous behavior. Arc represent that we don’t just want to judge through the outputs but look at the whole journey , the whole arc through interpretability.
We are interested more in the mechanisms than just output behaviors and then to create ai control and monitoring protocols on top of those. To prevent or control the above mentioned undesirable behaviors.
India is one country where AI adoption is rapidly growing and may surpass the global usage, furthering the need of initiatives such as ours to be present and work here and there are no similar major organizations most being based in the UK and the US.
Our Work:
Our current work spans Biosecurity, AI jailbreaks and emotion weight monitoring protocol. We are also simultaneously working on discovering new attack surfaces.
Some things that we have published are:
Loss Landscape Response to Adversarial Perturbation Is Architecture-Dependent
Conference Paper,Adversarial Robustness,TAIS
Toxin Feature Hierarchy in ESM-2
Workshop Paper,Protein LM,ICML GenBio
Fourier Gradient Regularisation for Adversarial Robustness
Workshop Poster,Adversarial Robustness,NeurIPS Reliable ML from Unreliable Data
More research we are exploring:
- Isolating Refusal and Compliance Heads: Unconditional Bipolar Defenses Against White-Box and Multi-Turn Attacks
- Thinking Model Emotions: Pre-Commitment and Functional Frustration in Extended Thinking
- The Geometry of Control: Spectral Attractors as Low-Dimensional Projections in Large Language Models
Community:
We are also trying to build a community for AI Safety around us, this includes giving talks in colleges around Delhi and maintaining and active Cohort where we take regular lectures and reading sessions around fundamentals of AI and AI Safety,
https://www.linkedin.com/company/redarc-labs/
Our goals:
Fellowships have played a meaningful role in our journey. Most of the opportunities we’ve seen are based in London or Berkeley, and they're often difficult for students and early-career researchers from India to access.
One of our goals is to help bridge that gap by creating opportunities and mentorship for people who want to contribute to AI safety from here.
Our goal for the next 6 months is to discover more attack surfaces and adversarial settings and start on a tool for multi-agent adversarial setting, we also plan on publishing 5-6 novel research artifacts to be developed and explored further.
It might sound ambitious and builds from India and goes toe to toe with major AI Safety orgs like Redwood, Goodfire Grayswan, METR.
Our Team:
We are two dedicated researchers:
Shivam Dubey
- Apart Research fellow, under Jason Hoelscher-Obermaier.
- MARS V Research Fellow, Cambridge AI Safety Hub.
- Lead on FASD project (77% bias reduction), cited by MIT Technology Review.
- GitHub: github.com/punctualprocrastinator · LinkedIn: linkedin.com/in/syntaxsavant · shivam@redarclabs.com
Manan Wadhwa
- MARS V Research Fellow, Cambridge AI Safety Hub
- Google Summer of Code 2026, HumanAI organisation.
- Research Fellow, AISI @ Georgia Tech.
- GitHub: github.com/Manan-Wadhwa · LinkedIn: linkedin.com/in/manan-wadhwa · manan@redarclabs.com
We are currently doing internships side by side to build credibility and to self fund some parts of the research and conducting sessions and workshops.
We also have small cohort of 6 fellows learning with us.
Theory of Impact
As AI systems become more capable, we think it is becoming harder to judge whether they are actually safe just by looking at their outputs. A model might refuse harmful requests during evaluation but behave very differently under adversarial prompting, long conversations, or when optimized for different objectives. Looking only at behavior tells us what happened, but often not why it happened.
We think that understanding the internal mechanisms behind a model's decisions will become increasingly important. If we can identify the circuits and representations that lead to behaviors like compliance, refusal, deception, or jailbreak susceptibility, we have a better chance of building safety interventions that remain useful even when models are pushed into adversarial settings.
This is the direction Redarc Labs is focused on. We study mechanistic interpretability under adversarial conditions. Most of our work starts with a simple question: what is happening inside the model when it succeeds or fails in a safety-critical setting? We then try to use those insights to build monitoring and control protocols instead of relying only on output-based evaluations.
Our current work looks at jailbreaks, biosecurity-related model behavior, emotion-like representations in reasoning models, and other potential attack surfaces that we think have not been explored enough. We are also interested in multi-agent settings because many future AI systems are likely to involve multiple models interacting with each other, creating new ways for failures to emerge.
Over time, we want this work to contribute to mechanism-based safety tools. If we can reliably detect when particular internal computations are occurring, those signals could eventually be used to monitor frontier models, improve evaluations, or build stronger control methods. We do not expect any single interpretability result to solve alignment, but we do think these tools become more valuable as models become more capable.
A second part of our work is building research capacity in India. There are excellent AI safety organizations in places like London and Berkeley, but relatively few opportunities for students here who want to work on mechanistic interpretability or frontier AI safety. We have benefited a lot from fellowships ourselves, and we want to create similar opportunities locally through reading groups, research cohorts, talks, and mentorship. If more researchers from India are able to contribute to frontier AI safety, we think that strengthens the field as a whole.
In the next six months, we want to explore new attack surfaces, publish several research artifacts, and begin building tools for adversarial multi-agent settings. Even if only some of these ideas work, we hope they help move safety research a little further away from treating models as black boxes and a little closer to understanding the computations that actually produce dangerous behavior.
https://docs.google.com/spreadsheets/d/1Xeiy4Ha9HXIqAJcKY-BE6GlMfTtcGnyGiTqqRDo0pOo/edit?usp=sharing
Our minimum budget of $20,000 would allow us to run Redarc Labs over the next six months while dedicating substantially more time to research. It would support our core researchers, provide the compute needed for mechanistic interpretability experiments, and help us continue building an AI safety community in India.
Our ideal budget of $40,000 would allow us to scale the lab rather than simply sustain it.
Additional funding would be used to:
- Expand compute for larger mechanistic interpretability and adversarial robustness experiments.
- Increase support for our research fellows so they can spend more time on independent research.
- Run additional workshops, reading groups, and technical events for students interested in AI safety.
- Support travel to AI safety conferences and workshops to present research and build collaborations.
- Invest in open-source tooling and research infrastructure that improves reproducibility.
- Bring on additional part-time research contributors as projects mature.
Our priority is to spend the majority of funding directly on research, researcher time, and the infrastructure needed to produce open, reproducible AI safety work. Community activities are included because they help develop the next generation of AI safety researchers in India and strengthen the local research ecosystem.