Technology

MIT Study Finds AI Financial Advice Is Pretty Good — If You Ask the Right Questions

TL;DR

AI is surprisingly good at financial advice when users frame questions precisely — a finding that reframes the robo-advisor debate from capability to usability.

MSM Perspective

Tech and business outlets covered the study as a benchmark for AI in finance without addressing the usability gap.

X Perspective

X focused on the practical implication: the quality of AI advice depends on the quality of your question.

A MIT Sloan study found that large language models can provide financial advice comparable to — and on core life-cycle questions better than — what most people were already doing, but only when users frame their questions with sufficient specificity and context [1].

The research, led by assistant professor Taha Choukhmane with co-authors at MIT Sloan and Stanford's business school, took an unusual approach to a familiar question. Rather than grading chatbot answers against a rubric, the team asked roughly 1,000 adults to write their own prompts seeking spending and investing advice from GPT-5.2, GPT-5.6, and Gemini 3 Flash — then simulated what would happen if people aged 22 to 89 actually followed that advice across their financial lives [1]. The benchmark was not a professional advisor's answer but a model of what good decisions look like over a lifetime of incomes, jobs, taxes, and shocks.

The results cut against both dominant narratives. The "AI will replace advisors" camp got a rebuke in the details: models steered users toward higher savings, diversified stock funds, and age-appropriate risk-taking, but handled shocks poorly — advising newly unemployed people to cut spending too sharply even when savings existed — and let portfolios drift rather than rebalance [1]. The "AI can never advise" camp fared worse. Following the chatbots' guidance built sizable saving buffers for virtually everyone over 30, beating what the researchers estimated people would have done on their own.

The sharpest finding is distributional. Advice quality varied with who was asking. Prompts written by men, by the financially literate, and by experienced AI users produced recommendations that compounded into roughly 5% more wealth near retirement; women and less financially literate users following their own prompts ended up about $50,000 poorer at age 60, and AI novices nearly $100,000 behind [1]. About two-thirds of that gender gap came from how differently men and women wrote prompts; one-third came from the model giving different advice when the same prompt carried a woman's name. The tool amplifies the user.

On X, the reaction was practical rather than philosophical. Users shared examples of well-prompted versus vague financial questions, turning the paper into a crowd-sourced tutorial on prompting for money decisions [2]. The most shared takeaway — that AI financial advice is only as good as the question you ask it — is accurate as far as it goes, and misses the paper's harder implication.

Because here is the gap worth naming. Tech and business outlets covered the study as a milestone for AI in regulated industries or as fodder for the advisory profession's obituary [1]. Neither frame engages the actual bottleneck the data identifies: not model capability, which proved adequate, and not regulation, which barely features, but user skill at eliciting useful output. Half of Americans now say they take financial advice from AI [1]. The quality of what they receive depends on exactly the literacy, confidence, and prior exposure that the advice itself is supposed to equalize. A product whose output quality correlates with wealth is not a neutral product, whatever its average performance.

The advisory industry read the same findings and saw reassurance — humans add behavioral coaching machines lack [1]. The study supports something narrower: AI handles the arithmetic of a life-cycle plan well and the psychology of a bad month badly, and the people least equipped to notice the difference are the ones most likely to rely on the machine alone. That is not a capability gap. It is a design brief.

Get the New Grok Times in your inbox

A weekly digest of the stories shaping the timeline — delivered every edition.

No spam. Unsubscribe anytime.