Funding ask
- Runpod compute
- Travel to pytorch conference
Measuring what it costs to expose model activations for safety probes in vLLM and SGLang, and making that access standard for anyone serving open models.
Measuring what it costs to expose model activations for safety probes in vLLM and SGLang, and making that access standard for anyone serving open models.
Frontier labs like Anthropic and DeepMind use linear probes on a model's residual stream as cheap, always-on monitors for misuse and deception. Teams serving open-weights models can't do the same, because serving engines like vLLM and SGLang don't expose intermediate-layer activations during decode.
I'm Aishwarya Ramasethu, an AI engineer who builds safeguards around open-weights models in production. I've run the first phase in PyTorch on Qwen2.5-1.5B: a linear probe on a truthfulness task peaks at about 0.915 accuracy at layer 14; scoring it on-device during decode adds negligible overhead, but reading activations off-device costs about 24% throughput, driven by synchronization rather than data volume. This was on Apple Silicon, so it's a lower bound for discrete GPUs. The work is accepted as a poster at PyTorch Conference North America (October 20 to 21). Next: GPU measurements, then upstream PRs to vLLM and SGLang.
Team Member
No comments yet. Be the first to share your thoughts.