U of A study identifies key AI weaknesses
A new study out of the University of Arizona compares some of the most commonly used AI tools to identify core weaknesses and which are the most fallible.
The study, published in Nature's Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.
Among the seven LLMs tested – ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5, Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 – they found that:
- ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements; Claude 3.5 Sonnet was the least.
- All seven were more susceptible to misinformation on obscure topics, implying that more training data on a given topic leads to more robust resistance to misinformation.
- DeepSeek was the most persuadable, as measured by responses to increasingly argumentative prompts, mostly because of its tendency toward sarcastic answers, which could not be reliably interpreted.
- Four models – ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek – corrected errors 100% of the time when given a second opportunity.
"This underscores the need for careful human engagement and the danger of blind reliance," said senior study author Dr. Marvin Slepian, Regents Professor of medicine and biomedical engineering. "When generative AI came out in November 2022, there was a lot of regulation potential, but that has since fell by the wayside. People are recognizing the onus is now left to the users."
Many people are familiar with AI's limitations, such as a tendency toward sycophancy, or the tendency to agree with users, and hallucination, or confidently wrong answers, but there has been very little work on evaluating AI's limitations during what is called multi-turn conversations, in which answers are predicated on previous context. Such usage more closely mirrors the real-world, according to the research team.
"These limitations raise important safety concerns, particularly as generative AI systems are increasingly deployed in high-stakes settings," Slepian said.