AI RESEARCH
Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories
arXiv CS.AI
•
ArXi:2606.04778v1 Announce Type: new Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shallow safety, where alignment concentrates in the first few output tokens. We show that shallow safety is a special case of a broader inference-time vulnerability, in which short token injections at any generation step can substantially alter subsequent safety behavior.