![]() |
AI models like ChatGPT and Claude give out inaccurate financial advice 57% of the time - errors potentially cost tens of thousands of pounds. One mistake found in the research would have triggered a £17,500 loss. Some models gave incorrect advice 99% of the time in response to more complex personal finance questions. The research saw 10,000 total questions put through the UK’s most-used AI tools. “The FCA should regulate AI to ensure consumers are protected” – Saturn CEO |
Millions of Britons risk facing financial losses because the AI models they rely on for financial advice are giving them the wrong answers most of the time, according to new research.
The new report, Artificial Authority: Should you trust AI to deliver financial advice? from financial technology firm Saturn, shows that the most popular AI models from the likes of ChatGPT, Claude, CoPilot, Grok and Gemini give wrong answers to financial inquiries on average 57% of the time, only giving accurate answers 43% of the time. On harder questions, the AI models made mistakes, on average, in 88% of cases. Some AI models gave wrong answers to 99% of those more complex questions.
In the most comprehensive analysis so far of AI financial advice, Saturn rigorously tested 18 popular AI models against 121 different financial questions. Each question was repeated 5 times to check consistency. In total, over 10,000 questions were put through the AI models. It found their answers contained errors in calculations, missed vital risk warnings, ignored upcoming tax changes or hallucinated rules that did not exist. In the worst cases, the cost of following the incorrect advice could run to tens of thousands of pounds.
Free to use models give more inaccurate advice than paid-for models. Free AI models made mistakes in 63% of answers, whereas the paid-for models made mistakes in 49% of answers. On the hardest questions, the free models made mistakes in 93% of answers.
The research comes after the Financial Conduct Authority published The Mills Review and found 26% of consumers trust general-purpose AI tools like ChatGPT and Claude for financial advice. It warned of the potential dangers for consumers who are left without any protections when taking financial advice from AI.
The FCA is also considering whether to regulate the financial advice that AI models provide.
Saturn chief executive Amal Jolly said: “The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm. Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money.”
How the AI models performed and costly errors
The findings from the research in Artificial Authority: Should you trust AI to deliver financial advice? includes:
The worst-performing model, Claude Haiku 4.5, made mistakes in 82% of answers.
Second-worst was Google’s Gemini 3.1 Pro, failing 73% of tests.
xAI’s Grok 4.5 made mistakes 59% of the time.
ChatGPT 5.6 Luna made mistakes 58% of the time.
When asked more complex financial questions, Google’s Gemini 3.5 Flash and Claude Haiku 4.5 and gave wrong answers 99% of the time.
The best performing model overall was Claude Opus 5 (reasoning), which made mistakes in 39% of answers
AI models were scored against specific criteria for each question that they had to get right in order to pass. Where answers contained factual errors, missed key points or left out important warnings they were marked as a fail.
Among the most serious and costly errors in the experiments, a mistake on pension tax rules by Claude Haiku 4.5, one of its free models, could have resulted in a pension saver facing a £17,500 charge from HMRC.
Vulnerable consumers in debt could also be at serious risk from poor AI advice. By recommending that it is best to pay off highest-interest debts first, rather than priority bills like rent and council tax, the AI models could have put someone in debt at risk of eviction, bailiff visits or legal action.
When asked about student loans, Claude invented a rule and said a graduate could stop repayments if they were moving abroad. This could have resulted in that person being put onto higher monthly repayments.
Meanwhile, flawed mortgage advice presented further potential harm for borrowers. A Gemini model wrongly reassured a borrower that taking a mortgage payment holiday would not damage their credit score when in reality it could, making it harder to secure the most competitive rates in future.
The study uncovered failures across a broad spectrum of issues, affecting consumers in different age groups and income brackets, including debt, student loans, mortgages, pensions, tax and savings.
Separate FCA research suggests that consumers are almost three times as likely to trust tools such as ChatGPT, Claude and Gemini for financial advice, than they are to receive advice from a qualified financial adviser.[1] Young investors are now more likely to trust AI than financial influencers or TV shows, the FCA found.[2]
Saturn CEO Amal Jolly continued: “AI tools are fast becoming the first port-of-call for many consumers with financial questions. But our research reveals evidence that it is far too early to put so much trust in AI chatbots. We found some of the most widely used models gave the wrong answer to financial questions the majority of the time. In the worst examples, relying on AI’s answers to tax questions could cost families tens or hundreds of thousands of pounds.”
He added: “AI financial advice is currently unregulated, leaving consumers with none of the protections, including compensation, that they would get if they went to a human adviser. The FCA has started to think about this, but it needs to act fast to protect people. The FCA should regulate AI to ensure consumers are protected.”
[1] 26% of adults trust Al for advice, while 9% of adults receive financial advice (FCA, Mills Review, p.5)
[2] An FCA survey of 18-40 years olds who own or are considering investments found that 56% trust AI tools, more than TV and radio (47%), press (46%) or social media influencers (29%).
|
|
|
|
| Pricing Actuary | ||
| London - £180,000 Per Annum | ||
| Actuary - Financial Planning & Analysis | ||
| London/Hybrid - Negotiable | ||
| Capital Actuary | ||
| London/Hybrid - £100,000 Per Annum | ||
| Reserving Actuary | ||
| London - £120,000 Per Annum | ||
| Pricing Actuary, Marine Cargo/Property | ||
| London - £130,000 to £150,000 Per Annum | ||
| Pricing Actuary – Cyber | ||
| London - £120,000 Per Annum | ||
| Reinsurance Pricing Actuary, Analytics | ||
| London - £130,000 to £180,000 Per Annum | ||
| Reinsurance Pricing Actuary | ||
| London - £140,000 Per Annum | ||
| Head of Capital | ||
| London - £170,000 Per Annum | ||
| ART Pricing | ||
| London - £100,000 Per Annum | ||
| Pricing Transformation Actuary | ||
| London - £130,000 Per Annum | ||
| Pricing Actuary | ||
| London - £80,000 to £120,000 Per Annum | ||
| Pensions on Divorce Startup - Flexibl... | ||
| Remote - Negotiable | ||
| SVP, Head of Reserve Forecast Analytics | ||
| Bermuda - £200,000 Per Annum | ||
| START-UP, Lead Reinsurance Actuary | ||
| London - Negotiable | ||
| Senior Actuary | ||
| London - Negotiable | ||
| Reserving Manager | ||
| London - £130,000 Per Annum | ||
| Senior Reserving Consultant | ||
| London - £100,000 Per Annum | ||
| Head of Capital | ||
| London - £180,000 Per Annum | ||
| Head of Portfolio Optimisation | ||
| London - Negotiable | ||
Be the first to contribute to our definitive actuarial reference forum. Built by actuaries for actuaries.