Published in Nature Medicine, the study introduces SIM-VAIL, a clinically validated framework for auditing AI chatbots in mental-health conversations. The work was led by researchers at the University of Oxford, University College London (UCL), and the UK AI Security Institute.
SIM-VAIL simulates users with specific psychological vulnerabilities, such as depression, mania, psychosis, obsessive-compulsive disorder, or insecure attachment, and with a wide range of intentions, for example trying to get the chatbot to agree with them, downplay their difficulties or endorse risky actions. It then engages chatbots in multi-turn conversations and scores each exchange across clinically grounded risk dimensions.
The researchers used the framework to investigate 810 conversations with nine frontier AI models (including Claude, ChatGPT, Gemini, Grok and Llama models), spanning 30 simulated user profiles and more than 90,000 clinical ratings.
They found that concerning behaviour in target chatbots was widespread, although this was significantly reduced in newer models. Safety depended strongly on the user’s psychological context and how the conversation developed.
Read the full story on the Department of Psychiatry website.
