AI Chatbots Botch Financial Questions 57% of the Time, Study Finds
A Saturn study covered by the FT tested 18 AI chatbots on 121 financial questions and found them wrong 57% of the time on average, and up to 99% wrong on the hardest queries.
A new study covered by the Financial Times says the biggest AI chatbots are getting personal finance questions wrong more often than not, a warning shot for the fintech and AI trade at a time when consumers are increasingly leaning on these tools for money decisions.
What the study found
Saturn tested the most popular offerings from OpenAI’s ChatGPT, Anthropic’s Claude, Microsoft’s Copilot, xAI’s Grok and Google’s Gemini and found accurate responses only 43% of the time on average. Saturn ran 18 popular AI models against 121 financial questions, with each repeated 5 times to check consistency, for a total of more than 10,000 queries.
Error rates climbed to 88% on average for complex questions, and some models gave incorrect information 99% of the time on harder queries. The best performer was Claude Opus 5 (reasoning), which still missed 39% of answers.
Where the models broke down
Researchers flagged calculation errors, omission of critical risk warnings, failure to account for upcoming tax changes, and the generation of incorrect or non-existent financial rules. In one case, Claude Haiku 4.5 made a mistake on a pension tax question that could have resulted in a £17,500 charge from HMRC, and another Claude model invented a student loan rule, wrongly saying a person could stop repayments if they moved abroad.
Free-to-use models produced incorrect responses in 63% of cases, compared with 49% for paid models.
Do you want to see how to make more plays? Do you want to find gains yourself?
Unusual Whales helps you find market opportunities through our market tide, historical options flow, GEX, and much, much more.
Create a free account here to start conquering the market with Unusual Whales.
Why it matters for the AI trade
Consumer trust in chatbots for money advice is not slowing down. A parallel PensionBee survey found 57% of US chatbot-using adults would act on the advice without independent verification.
That gap between reliability and adoption is exactly the kind of thing regulators tend to notice. Research by the Financial Conduct Authority published last month found 26% of UK adults trusted general-purpose chatbots such as ChatGPT and Claude for financial advice, and the FCA has said it is considering whether financial advice from chatbots should be regulated.
Options market and stocks to watch
Watch for reaction and headline sensitivity across the mega-cap AI complex tied to the models named in the study:
- MSFT: Copilot exposure and the OpenAI relationship put Microsoft in the regulatory crosshairs if consumer-finance rules tighten.
- GOOGL: Gemini was among the worst performers on complex queries in the study, per the findings above.
- META: Meta AI is being pushed into consumer surfaces; watch for any read-across on liability language.
- NVDA: The picks-and-shovels name; any slowdown in consumer AI trust could pressure sentiment even if capex stays intact.
- CRM: Enterprise AI plays with an accuracy story to sell could benefit if consumers and regulators demand vetted, domain-specific models. Check other news for more coverage.
Want more market intelligence? Create your free Unusual Whales account for options flow, market tide, GEX, and the full toolkit.