RESEARCH PROJECT · 2026

SafeRI

Recognition and Intervention for Token-Level
Safety Intervention in Large Vision Language Models

Caoyuan Ma2,3,* · Tian Gu1,5,* · Wenpu Liu3,4 · Weichu Xie3,4 · Shuai Dong3,6 · Yuqi Xu3,4 · Ji Zhao3 · Ziyue Wang3,4 · Wenzheng Chang3,7 · Taiqiang Wu3,8 · Yongfu Zhu3 · Wenqi Shao3,† · Zheng Wang1,‡ · Yinqiang Zheng2,‡

1 Wuhan University · 2 The University of Tokyo · 3 JD.com · 4 Peking University
5 Shanghai AI Laboratory · 6 Shanghai Innovation Institute · 7 Shanghai Jiao Tong University · 8 The University of Hong Kong

* Equal contribution   † Project leader   ‡ Co-corresponding authors

ABSTRACT

Safety alignment should be an on-demand intervention, not a permanent modification to every decoding trajectory.

Existing safety alignment methods modify model behavior globally, so their safety parameters affect both unsafe and already-safe generations. SafeRI instead combines streaming recognition with gated LoRA intervention: a lightweight recognizer monitors each pre-token generation state and activates the safety adapter only when unsafe drift emerges. The adapter learns from unsafe prefixes, transition statements, and safe continuations to redirect risky generations, while safe trajectories retain the frozen-backbone policy. Experiments across safety and general-purpose benchmarks show effective post-alignment hardening with limited impact on multimodal utility.

01 / METHOD

Recognize risk.
Intervene only then.

01

Frozen VLM

Generate normally and retain the original policy on safe trajectories.

→
02

Recognizer

Read the current prefix state and estimate imminent unsafe drift.

→
03

Intervention

Open a renewable intervention window and steer the continuation safe.

Figure 1. SafeRI performs streaming recognition at every decoding step. The safety adapter remains dormant until risk emerges.

Token-level

The gate is updated from the current hidden state for the next decoding step, making intervention causal rather than retroactive.

Selective

Once risk subsides, the gate closes and generation returns to the frozen-backbone policy.

02 / RESULTS

Safety gains,
utility retained.

SafeRI improves the safety average on all evaluated backbones while keeping general multimodal capability close to the frozen base model.

BackboneSettingSafety Avg. ↑General Avg. ↑Safety Δ
Qwen3.5-9BBase86.5667.85—
Qwen3.5-9BSafeRI88.8867.64+2.32
Qwen3.5-4BBase87.5766.32—
Qwen3.5-4BSafeRI88.2864.85+0.71
Qwen3.5-2BBase86.5461.05—
Qwen3.5-2BSafeRI87.7062.21+1.16
Llama3.2-Vision-11BBase83.4959.45—
Llama3.2-Vision-11BSafeRI84.2758.93+0.78

Safety Avg. is the arithmetic mean over SPA-VL-test harm, AdvBench, HADES, XSTest, and MSSBench. Higher is safer.

SPA-VLAdvBenchHADESXSTestMSSBenchMMBenchMM-VetBLINK

03 / CONTRIBUTIONS

01We formulate post-alignment VLM safety hardening as a selective trajectory-control problem.

02We introduce streaming hidden-state recognition paired with a boundary-aligned LoRA intervention.

03We show that selective activation can reduce residual safety risk without globally shifting safe behavior.

CITATION

Read and cite
the paper.

The complete SafeRI preprint is now available on arXiv.

Download PDF →
@misc{ma2026saferi,
  title        = {SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models},
  author       = {Caoyuan Ma and Tian Gu and Wenpu Liu and Weichu Xie and Shuai Dong and Yuqi Xu and Ji Zhao and Ziyue Wang and Wenzheng Chang and Taiqiang Wu and Yongfu Zhu and Wenqi Shao and Zheng Wang and Yinqiang Zheng},
  year         = {2026},
  eprint       = {2609.03544},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url          = {https://arxiv.org/abs/2609.03544}
}