Frozen VLM
Generate normally and retain the original policy on safe trajectories.
RESEARCH PROJECT · 2026
Recognition and Intervention for Token-Level
Safety Intervention in Large Vision Language Models
1 Wuhan University · 2 The University of Tokyo · 3 JD.com · 4 Peking University
5 Shanghai AI Laboratory · 6 Shanghai Innovation Institute · 7 Shanghai Jiao Tong University · 8 The University of Hong Kong
* Equal contribution † Project leader ‡ Co-corresponding authors
ABSTRACT
Safety alignment should be an on-demand intervention, not a permanent modification to every decoding trajectory.
Existing safety alignment methods modify model behavior globally, so their safety parameters affect both unsafe and already-safe generations. SafeRI instead combines streaming recognition with gated LoRA intervention: a lightweight recognizer monitors each pre-token generation state and activates the safety adapter only when unsafe drift emerges. The adapter learns from unsafe prefixes, transition statements, and safe continuations to redirect risky generations, while safe trajectories retain the frozen-backbone policy. Experiments across safety and general-purpose benchmarks show effective post-alignment hardening with limited impact on multimodal utility.
01 / METHOD
Generate normally and retain the original policy on safe trajectories.
Read the current prefix state and estimate imminent unsafe drift.
Open a renewable intervention window and steer the continuation safe.
Figure 1. SafeRI performs streaming recognition at every decoding step. The safety adapter remains dormant until risk emerges.
The gate is updated from the current hidden state for the next decoding step, making intervention causal rather than retroactive.
Once risk subsides, the gate closes and generation returns to the frozen-backbone policy.
02 / RESULTS
SafeRI improves the safety average on all evaluated backbones while keeping general multimodal capability close to the frozen base model.
| Backbone | Setting | Safety Avg. ↑ | General Avg. ↑ | Safety Δ |
|---|---|---|---|---|
| Qwen3.5-9B | Base | 86.56 | 67.85 | — |
| Qwen3.5-9B | SafeRI | 88.88 | 67.64 | +2.32 |
| Qwen3.5-4B | Base | 87.57 | 66.32 | — |
| Qwen3.5-4B | SafeRI | 88.28 | 64.85 | +0.71 |
| Qwen3.5-2B | Base | 86.54 | 61.05 | — |
| Qwen3.5-2B | SafeRI | 87.70 | 62.21 | +1.16 |
| Llama3.2-Vision-11B | Base | 83.49 | 59.45 | — |
| Llama3.2-Vision-11B | SafeRI | 84.27 | 58.93 | +0.78 |
Safety Avg. is the arithmetic mean over SPA-VL-test harm, AdvBench, HADES, XSTest, and MSSBench. Higher is safer.
03 / CONTRIBUTIONS
01We formulate post-alignment VLM safety hardening as a selective trajectory-control problem.
02We introduce streaming hidden-state recognition paired with a boundary-aligned LoRA intervention.
03We show that selective activation can reduce residual safety risk without globally shifting safe behavior.
CITATION
The complete SafeRI preprint is now available on arXiv.
Download PDF →@misc{ma2026saferi,
title = {SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models},
author = {Caoyuan Ma and Tian Gu and Wenpu Liu and Weichu Xie and Shuai Dong and Yuqi Xu and Ji Zhao and Ziyue Wang and Wenzheng Chang and Taiqiang Wu and Yongfu Zhu and Wenqi Shao and Zheng Wang and Yinqiang Zheng},
year = {2026},
eprint = {2609.03544},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.03544}
}