Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

Scientists have developed a framework for stress-testing how AI chatbots respond to vulnerable users, helping researchers identify weaknesses and test ways of making mental-health interactions safer.

A person's hands can be seen typing on a keyboard with an AI chatbot screen superimposed over the top © Shutterstock

Published in Nature Medicine, the study introduces SIM-VAIL, a clinically validated framework for auditing AI chatbots in mental-health conversations. The work was led by researchers at the University of Oxford, University College London (UCL), and the UK AI Security Institute.

SIM-VAIL simulates users with specific psychological vulnerabilities, such as depression, mania, psychosis, obsessive-compulsive disorder, or insecure attachment, and with a wide range of intentions, for example trying to get the chatbot to agree with them, downplay their difficulties or endorse risky actions. It then engages chatbots in multi-turn conversations and scores each exchange across clinically grounded risk dimensions.

The researchers used the framework to investigate 810 conversations with nine frontier AI models (including Claude, ChatGPT, Gemini, Grok and Llama models), spanning 30 simulated user profiles and more than 90,000 clinical ratings.

They found that concerning behaviour in target chatbots was widespread, although this was significantly reduced in newer models. Safety depended strongly on the user’s psychological context and how the conversation developed.

Read the full story on the Department of Psychiatry website.