> Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
[AUTHOR: Tony Li, Hongliang Liu and Yuhao Wu]
[DATE: 28/08/2026 22:00]
[LANGUAGE: EN]
New research reveals that AI safety refusal lives in a thin neural layer, highlighting the critical need for external, multi-layered security.
The post Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety appeared first on Unit 42.