A project about researching radical methods to both advance and prepare countermeasures for what is to come after transformers.
A project about researching radical methods to both advance and prepare countermeasures for what is to come after transformers.
Project Details
Updated 07/24/26 · Provided via application · VerifiedCurrent research on reducing AI risks assumes that future frameworks are going to be similar to today’s Transformer-driven Large Language Models, despite the rapid progress of AI research. This assumption might not necessarily be true. There are many constraints to consider for Transformer models. As current laboratories work towards improving efficiency, longer context, and stronger reasoning, other architectures may become more and more relevant, including state-space models, recurrent architectures, and linear-attention mechanisms, as well as novel and creative structures that are yet to be discovered.
The goal of this project is to explore architectures like Mamba, RWKV-7 "Goose", Gated DeltaNet, or frameworks building upon them, such as KDA (Kimi Delta Attention) in hybrid and non-hybrid systems. I am not just trying to replicate these architectures, but to further explore their limitations, benefits, and whether a truly unique design might be possible, and ways to implement to tackle associated risks.
Why me? - I have been researching mathematics and machine learning as an independent researcher for a couple of years, solely funded by my savings and income. In the past year. I have studied and experimented with self-attention transformers, steady-state models such as Mamba 2 and Mamba 3 (SISO/MIMO), and even Linear attention frameworks like Gated Deltanet and its versions such as KDA (Kimi Delta Attention) at both the mathematical level (I do not claim to have contributed to the research papers of any of these frameworks) and experiments with those frameworks and how they work not only at training, but also at inference.
Theory of Impact
Updated 07/24/26 · By grantmaking.aiCurrent AI risk mitigation research primarily targets the current dominant mode of AI, and they rarely consider the frameworks that may be dominant in the future. By researching and evaluating the capabilities of frameworks that may define the future, I aim to prepare for those risks in advance.
People
Updated 07/24/26 · By grantmaking.aiTeam Member
I do not know the applicant personally and cannot vouch for their track record, but the premise of the proposal is sound: AI safety research remains heavily anchored to Transformer architectures, leaving a real gap around the state-space, recurrent, and linear-attention alternatives that may come to dominate. Work that gets ahead of that shift, rather than reacting to it after the fact, strikes me as both timely and worth supporting, and I endorse this proposal on the strength of that direction.