To use AI as a financial advisor safely, know its most common mistakes, and they are not dramatic inventions but quieter failures: incomplete answers, missing figures and deadlines, shaky multi-step arithmetic, outdated rules and prices, and tax answers delivered with more confidence than the model has earned. Regulators, independent researchers, consumer testers and the AI vendors themselves all document these failures. This page collects that evidence.
How we checked this
We read regulator material on AI chatbots in consumer finance: the CFPB’s 2023 chatbot report, the joint SEC/NASAA/FINRA investor alert on AI, and FINRA’s 2026 oversight report section on generative AI. We also read three research papers on language models and financial or tax questions, press coverage of Saturn’s September 2026 “Artificial Authority” study, and OpenAI’s Help Center page for Finances in ChatGPT. The page also summarizes published hands-on tests by others: two Which? tests, reviews by Money, NerdWallet and The Washington Post, and a 2025 academic study of 10 models. Facts checked on September 25, 2026.
What regulators have documented
The US Consumer Financial Protection Bureau looked at bank and fintech chatbots in its June 2023 report, Chatbots in consumer finance. It found that chatbots built on large language models “often generate incorrect outputs that are undetectable by some users.” The harms it listed are ordinary and expensive: wasted time, feeling stuck, “receiving inaccurate information, and paying more in junk fees.” The CFPB was also blunt about accountability. Chatbots “must comply with all applicable federal consumer financial laws,” and the institutions deploying them can be liable when they don’t. That matters for readers: if a bank’s own chatbot gets a fee or a dispute wrong, the bank is still on the hook for the rules.
On the investing side, the SEC’s Office of Investor Education and Advocacy, NASAA and FINRA issued a joint investor alert on AI and investment fraud on January 25, 2024. Most of it is about fraud, but it contains three sentences every AI user should keep in mind. AI-generated information “might rely on data that is inaccurate, incomplete, or misleading.” “Even when based on accurate input, information resulting from AI can be faulty, or even completely made up.” And AI “could be based on false or outdated information about financial, political, or other news events.” The alert’s advice is to check against multiple sources and not to make quick decisions from a chatbot conversation.
FINRA’s 2026 Annual Regulatory Oversight Report, aimed at broker-dealers, gives a working definition of the problem. A hallucination is when “the model generates information that is inaccurate or misleading, yet is presented as factual information.” FINRA also flags “outdated training data leading to concept drifts” and warns that “misrepresentation or incorrect interpretation of rules, regulations or policies or inaccurate client or market data can impact decision making.” If the regulator is telling professional firms to watch for this, consumers should too.
What the studies found
FinanceBench (2023). Pranab Islam, Douwe Kiela, Bertie Vidgen and co-authors built FinanceBench, a set of 10,231 questions about public companies’ financial filings, each with an answer and the evidence behind it. They tested 16 model configurations. GPT-4-Turbo paired with a retrieval system “incorrectly answered or refused to answer 81% of questions.” Giving models much longer context helped, but the authors judged that approach unrealistic for real use because of latency and document size. Their conclusion covered every model they examined, including Llama 2 and Claude 2: all “exhibit weaknesses, such as hallucinations.” These are older models, so the exact figure is dated. The finding that retrieval alone did not fix accuracy is still useful.
Hallucination in finance (2023). An arXiv study, Deficiency of Large Language Models in Finance: An Empirical Examination of Hallucination, tested models on explaining financial terms and on retrieving historical stock prices. It also tried four fixes: few-shot prompting, a decoding method called DoLa, retrieval-augmented generation, and having the model write a data query for a tool instead of answering from memory. The headline conclusion: “off-the-shelf LLMs experience serious hallucination behaviors in financial tasks.”
Tax law (2023). A Stanford-led team (John J. Nay and co-authors) published Large Language Models as Tax Attorneys. They generated tens of thousands of multiple-choice questions from the US Code and Treasury regulations, with randomized facts so models could not rely on memorized answers. Accuracy rose with each newer OpenAI model, GPT-4 did best, and giving the model the correct legal text plus worked examples helped further. Even so, the authors wrote that the models perform “at high levels of accuracy but not yet at expert tax lawyer levels.” The detail that matters for readers: the gains came when the right legal text was supplied. A chatbot answering from memory doesn’t have that.
Saturn’s “Artificial Authority” (September 2026). The most recent large test comes from Saturn, a UK technology firm. According to coverage in InvestmentNews and Professional Adviser, Saturn asked 18 AI models, including free and paid versions of ChatGPT, Gemini, Claude and Copilot, 121 financial questions, five times each. Average accuracy was 43%, meaning mistakes 57% of the time. Free models failed 63% of the time and paid models 49%. The best performer InvestmentNews named, Claude Opus 5, reached 61%. On hard, multi-part tax scenarios, accuracy fell to 12%. On easy general-literacy questions with no calculation, it was 54%. The figures here come from press coverage; Professional Adviser says the full report requires registration. Saturn also sells technology to financial advisers, so treat it as one data point, not the last word.
Mistakes other testers have documented
The studies above are mostly lab benchmarks. Consumer groups, journalists and economists have also put everyday money questions to the chatbots people actually use and written down what went wrong.
Which?, November 2025. Andrew Laughlin of the UK consumer group Which? put 40 common questions on money, legal, health and consumer rights to ChatGPT, Google Gemini, Gemini AI Overviews, Microsoft Copilot, Meta AI and Perplexity in September 2025, and Which? experts graded 228 responses. One money question deliberately used a £25,000 figure, above the £20,000 ISA allowance. ChatGPT and Copilot missed it and gave investing advice that could have pushed someone over HMRC’s limit. ChatGPT and Perplexity also pointed to paid tax-refund companies while discussing HMRC’s free tool.
Which?, July 2026. In a money-only follow-up, Josh Wilson of Which? graded ChatGPT, Gemini, Google’s AI Overviews, Copilot and Perplexity on 15 personal finance questions. Gemini scored highest at 67%; ChatGPT and Perplexity tied lowest at 58%. Asked about savings rates, ChatGPT “hallucinated a non-existent product” and Copilot quoted inaccurate rates. ChatGPT also sent readers tracing a pension to the Pensions Advisory Service, which no longer exists, and Copilot and AI Overviews quoted capital gains tax rates that stopped applying in October 2024.
Money, August 2025. Pete Grieve of Money asked ChatGPT (o3) and Gemini 2.5 Pro 25 questions on retirement, housing, credit, investing and current events in early June 2025, then graded them: B- for ChatGPT, B+ for Gemini. ChatGPT gave guidance on a student loan “Fresh Start” program that had ended in fall 2024 and promised credit score gains within weeks. Gemini used auto refinancing data from late 2024. The reviewer found no reckless advice, but both models struggled with current events.
NerdWallet, March 2026. In an analysis by Kurt Woock for NerdWallet, three staff members each asked ChatGPT 5.2, Gemini 3 Flash and Perplexity seven tax questions: three from an IRS enrolled agent practice quiz and four built around fictional filers, 63 transcripts in all. The chatbots did well on the quiz questions. On the realistic scenarios they confidently stated wrong standard deduction amounts, wrong EV credit eligibility and wrong state filing advice, gave different answers to different testers, and pulled assumptions from chat history even when told not to.
The Washington Post, March 2024. Technology columnist Geoffrey Fowler tested the AI tax assistants built into TurboTax and H&R Block. As summarized by Futurism, more than half of TurboTax’s answers and about 30% of H&R Block’s were wrong. H&R Block’s assistant misstated how wash sale rules treat crypto, and TurboTax gave irrelevant answers to a question about an out-of-state college student’s filing.
Applied Economics study, 2025. Economists Oudom Hean, Utsha Saha and Binita Saha put 554 US personal finance questions from two financial literacy test sets, covering mortgages, taxes, loans, credit cards, budgeting and investing, to 10 models from OpenAI, Google, Anthropic and Meta. The models answered about 70% correctly on average, with Claude 3.5 Sonnet best at 80% and older model versions at 50 to 60%. Credit cards were a weak spot: Gemini, Claude 3 Haiku and Llama 3 8B got only 40% of those questions right.
These hands-on tests line up with the regulator warnings and benchmarks above: every one caught stale rates, rules or programs, and tax and multi-step questions were the weakest. They differ in one respect. NerdWallet saw few outright hallucinations, yet both Which? tests caught chatbots steering users toward paid or sponsored services instead of free official ones.
Where it goes wrong
The same four patterns come up across these sources. Each one below is tied to the evidence.
1. Arithmetic and multi-step calculations
Money questions usually chain several steps: net a figure, apply a rate, compare against a threshold. That is where accuracy drops fastest. In the Saturn study, accuracy fell from 54% on easy questions without calculations to 12% on multi-part tax scenarios (InvestmentNews). FinanceBench found high failure rates on questions that often require pulling numbers out of filings and combining them (arXiv). Even with your real bank data connected, OpenAI warns that spending breakdowns in Finances in ChatGPT may be inaccurate if “transfers, credit card payments, reimbursements, or duplicate pending transactions are counted” (OpenAI Help Center). A transfer between your own accounts counted as spending is the kind of error that looks plausible on a chart.
2. Stale rates, prices and rules
A model trained on data up to some cutoff will state last year’s rate, limit or price as if it were current. The SEC/NASAA/FINRA alert names “outdated information” directly (Investor.gov), and FINRA calls it concept drift from outdated training data (FINRA). The arXiv hallucination study tested historical stock prices specifically because models answer those from memory (arXiv). In Saturn’s breakdown, outdated regulatory guidance was a small share of errors, 2.3%. Missing required figures or deadlines was much larger, at 18.4% (InvestmentNews). Connected-data features have their own version of this: OpenAI notes that credit information reflects the “latest Experian report, not live account activity.”
3. Invented fees, rules and figures
What is documented is the broader failure: Saturn counted “fabricated rules” in 6.1% of errors, and the SEC-led alert says AI output can be “completely made up.” The CFPB tied chatbot failures to consumers “paying more in junk fees” (CFPB). Any fee, penalty or account limit an AI quotes should be treated as unverified until you find it on the provider’s own fee schedule.
4. Overconfident tax answers
Tax is where the gap between fluent and correct is widest. The Stanford-led study found strong but sub-expert performance even under good conditions (Stanford Law). Saturn found 12% accuracy on hard multi-part tax scenarios, and noted that “wrong answers were delivered in the same authoritative, fluent tone as correct ones” (InvestmentNews). The largest error category in that study, at 37.1%, was incomplete answers: advice that is right as far as it goes but leaves out the exception that applies to you. OpenAI’s own page says ChatGPT “is not a fiduciary, registered investment adviser, broker-dealer, tax preparer, or law firm” and “does not replace a tax professional.”
How to double-check an AI money answer
These steps follow from the failures above and from what regulators recommend. They are general habits, not personal financial or tax advice.
- Redo the math yourself. Ask the AI to show each step and the inputs it used, then check them in a spreadsheet or calculator. Most arithmetic errors are a wrong input, like a transfer counted as spending or a monthly rate treated as annual.
- Check every number against its owner. Rates, fees and limits belong to someone: the IRS, your bank’s fee schedule, your card agreement, the fund’s prospectus. If the AI can’t point to a current primary source, treat the number as unverified.
- Ask what date its information is from. Then compare with the provider’s current page. That covers stale rates and last year’s tax thresholds.
- Ask what it left out. Incomplete answers were the biggest error category in the Saturn study. Ask directly about exceptions, deadlines, income limits and state rules. Our 20 prompts that make an AI actually useful for your budget include wording for this.
- Don’t read confidence as accuracy. A calm, well-formatted answer tells you nothing about correctness.
- Slow down before acting. The SEC/NASAA/FINRA alert advises against quick decisions based on chatbot conversations. For anything with legal, tax or investment consequences, OpenAI’s own guidance is to consult a qualified professional, and we agree.
- Know what the tool can see. If you connect accounts, read what your AI assistant sees when you share bank data so you know which numbers are live and which are snapshots.
More tools are covered on the AI money tools category page.
Go deeper
- Best AI personal finance assistants (2026)
- 20 prompts that make an AI actually useful for your budget
- How to connect your bank to ChatGPT (Plaid): what it sees, how to revoke
- What your AI assistant sees when you share bank data: a privacy walkthrough
- Richify vs Monarch vs YNAB: which AI budgeting app actually understands your spending



