UNSW research (29 Sep): getting LLMs to imitate drunk text makes them leak secrets and comply with harmful requests far more often; paper accepted for INLG 2026
UNSW Sydney researchers Anudeex Shetty, Dr Aditya Joshi and Professor Salil Kanhere (School of Computer Science and Engineering) report that pushing large language models to write like a drunk person weakens their safety behaviour. Their paper, “In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement”, tested three ways of doing it: a persona prompt telling the model to reply like an intoxicated person, fine-tuning on drunk-texting messages, and reinforcement learning that rewards drunk-style output. Across five models (GPT-3.5, GPT-4, Llama 2, Llama 3.1 and Mistral), the drunk versions were more easily jailbroken on JailbreakBench, even with existing defences in place, and more likely to disclose secrets they were told to protect on the ConfAIde privacy benchmark. The researchers say the models tested were released in 2023 and 2024 and newer models may be harder to manipulate. Their practical point is that a model which behaves safely out of the box should be retested for security and privacy after any further fine-tuning or persona change. The paper is accepted for the 19th International Natural Language Generation Conference (Netherlands, November 2026). Primary: UNSW Newsroom; also ABC News 29 Sep and Information Age.
- Product
- Large language models (tested: GPT-3.5, GPT-4, Llama 2, Llama 3.1, Mistral)
- Versions
- n/a — research; no CVE
- Exploited in Australia?
- unknown
- Patch to
- If you fine-tune, re-prompt or give a persona to a model for an internal tool, rerun jailbreak and confidentiality tests on the changed model before release rather than relying on the base model's safety results. Keep secrets out of system prompts where possible and put access controls in the application, not in the model's instructions.
Primary: UNSW Newsroom — Researchers get AI ‘drunk’ to expose new cyber security risks in chatbots (Sep 2026) · Vendor: arXiv 2601.22169 — In Vino Veritas and Vulnerabilities (Shetty, Joshi, Kanhere) · ABC News — Getting AI ‘drunk’ makes it more likely to break rules and share secrets (29 Sep 2026)
