| |
 | | Securing AI Models Against Backdoors, Jailbreaks, and Prompt Injections |  | NAY Myat Min PhD Candidate School of Computing and Information Systems Singapore Management University FULL PROFILE |
Research Area - Information Systems & Technology
Dissertation Committee | Advisor: | | | Members: | | | | | | External Members: | ZHANG Tianwei, Associate Professor & Provost’s Chair in Computing, College of Computing and Data Science, Nanyang Technological University |
| | Date 3 August 2026 (Monday) Time 4:00pm – 5:00pm Venue Meeting room 5.1, Level 5 School of Computing and Information Systems 1, Singapore Management University, 80 Stamford Road, Singapore 178902 Please register by 2 August 2026. We look forward to seeing you at this research seminar.
|
| ABOUT THE TALK Deep neural networks (DNNs) and large language models (LLMs) face threats across training and deployment: backdoors, jailbreaks, prompt injections, and concept-conditioned semantic inconsistencies. Existing defenses often address these threats separately. This dissertation develops a signal-guided design principle: a matched perturbation, probe, or operating condition exposes a method-specific diagnostic signal that guides repair, audit triage, or abstention. This signal–perturbation–action structure unifies the logic without claiming a universal signal or failure mechanism. Four methods instantiate this principle. ULRL uses clean-data loss ascent to expose unusually malleable classifier rows, then reinitializes and relearns them. With 1% clean data, it reduces attack success to single digits in most cases across twelve attacks, four datasets, and three architectures, at a one-to-two-point accuracy cost. CROW regularizes layerwise consistency under adversarial embedding perturbations, mitigating six evaluated backdoors across five Llama-2, CodeLlama, and Mistral models using 100 clean samples but no trusted clean reference model. RAVEN provides black-box triage for concept-conditioned anomalies by combining response uniformity under paraphrase-controlled queries with divergence from selected peer models; it identifies controlled implants and flags high-suspicion patterns in nine of twelve topics across five pretrained LLMs. LCF monitors hidden-representation changes during prefill and abstains on anomalous prompts. With deployment-matched clean calibration, it reduces mean backdoor attack success below 1% on Qwen2.5-7B and Gemma-2, detects 92–100% of DAN-style jailbreaks, and flags all evaluated text-payload prompt injections at under 0.1% measured inference overhead. Across the studied threat models, these results support signal-guided repair, audit triage, and runtime monitoring. | ABOUT THE SPEAKER NAY Myat Min is a final-year PhD candidate in Computer Science at SMU’s School of Computing and Information Systems, supervised by Professor SUN Jun. His research interests lie in AI security and trustworthy machine learning, with a particular focus on protecting vision models and large language models against backdoors, jailbreaks, prompt injections, and other adversarial threats. His research has been published at ICLR 2026 and ICML 2025, and in IEEE Transactions on Information Forensics and Security. He also gained applied AI-safety research experience through an internship at the Singapore Digital Trust Mechanism Lab, Huawei Singapore Research Center. Nay holds an MSc in Cyber Security from Mahidol University and a BSc with First Class Honours in Computer and Network Technology from Northumbria University. In his leisure time, he enjoys swimming and playing chess. |
|