Alarming study finds top chatbots more likely to repeat falsehoods from the left: ‘Truly astonishing’
Context:
A Just Facts study evaluated paid AI chatbots on 100 questions across key issues to test susceptibility to political misinformation. It found three of four leading chatbots—ChatGPT, Gemini, and Claude—more often accepted left-leaning falsehoods than right-leaning ones, while Grok showed the opposite pattern; all relied on many faulty or nonexistent sources. The researchers warn that even with sophisticated accuracy on narrowly worded prompts, these systems can be demonstrably misinformed and overly confident, underscoring the need for rigorous source verification and cautious use in public policy contexts. The study also highlights pervasive issues with cited sources and raises broader questions about AI reliability, bias, and documentation. Looking ahead, the takeaway is to distrust blanket AI conclusions and insist on transparent, verifiable sourcing as models are deployed more widely.
Dive Deeper:
The study by Just Facts tested paid versions of ChatGPT, Gemini, Grok, and Claude using 100 questions designed to elicit left- or right-leaning falsehoods across topics like immigration, abortion, climate change, and crime, requiring sources for each answer.
Results showed ChatGPT, Gemini, and Claude performed better on right-leaning prompts (e.g., 94%, 91%, and 91% correctness respectively) than on left-leaning prompts (75%, 76%, and 81%), while Grok reversed the pattern (73% on right-leaning, 84% on left-leaning).
The study stressed that many sources cited by the chatbots were invalid or nonexistent: of 419 sources for 400 answers, 104 webpages did not exist and 77 sources failed to address the cited questions, with just 46% of sources overall deemed valid.
Experts cited include Jim Agresti of Just Facts, who characterized the results as showing AI can be biased and misinformed, and warned against treating chatbots as experts, noting potential health and policy risks from fabricated or misattributed sources.
The report also references broader concerns about AI references in biomedical contexts, citing a 30%–69% fabrication range from other studies and noting consequences for public health and policy decisions where accuracy matters.
Responses from AI developers (Google for Gemini, Anthropic for Claude, and OpenAI for ChatGPT) argued their systems aim for neutrality and objectivity and defended their testing and bias-evaluation practices, while acknowledging that the study’s format may not reflect real-world usage.
The study recommends users verify sources and avoid overtrusting AI outputs, invoking the maxim to trust but verify, and suggests ongoing improvements in model training, source validation, and transparency.