AI Financial Advice Gets Better When Prompts Get Specific
You know the moment: someone mentions using a chatbot for retirement planning, and the next thing you hear is a confident-sounding answer about stocks, budgets, and “optimal” saving. It feels helpful… until it doesn’t.
What’s interesting about the recent research on AI financial advice is that it points to a surprisingly practical reason for the gap between “sounds right” and “works right”: the quality of the questions matters almost as much as the model.
Large language models (LLMs) can produce advice that lines up with standard life-cycle financial planning ideas—save during working years, shift from risky assets to safer ones later, and build a meaningful buffer before relying on investments. But they can also miss key real-world nuances, especially when someone’s life throws a shock like unemployment. And when prompts are vague or incomplete, the model often falls back on rules of thumb.
A new MIT Sloan research effort digs into exactly that prompt-quality effect by running careful simulations over a lifetime.
The heart of the problem: LLM advice is “question-shaped”
An LLM (large language model) is a machine-learning system trained to predict what text comes next. When you ask a question, the model doesn’t just “know finance”; it generates an answer that fits the pattern of your prompt.
That matters because financial advice isn’t generic. Two people can face the same broad goal (“invest for retirement”) while needing completely different tactics based on details like:
- age and time horizon
- employment status
- income level
- existing savings and account balances
- how taxes and retirement benefits interact
So the question becomes: what makes the model’s output more “planner-like” and less “chatty?” In this research, the difference came from prompt structure.
In one version of the study, participants (1,000 adults) wrote their own prompts asking for spending and investing advice to an LLM. In another version, the prompts were replaced with “academic” prompts that included full financial information plus clear assumptions about the economic environment and retirement rules.
That setup makes the experiments feel a bit like comparing two pilots:
- one flight plan was handwritten with missing waypoints,
- the other included coordinates and constraints.
The study in plain language: simulating a whole life
To test whether advice is actually good over time, the researchers didn’t just judge the text. They turned the advice into a policy and simulated how a person’s financial life might evolve if they followed it.
Here’s the core pipeline, conceptually:
- Collect prompts: people asked for spending and investing guidance.
- Get AI advice: the LLM responded with recommendations.
- Translate recommendations into decisions: the advice was converted into measurable actions like saving rates and stock allocations.
- Simulate years forward: the researchers ran the clock from early adulthood through retirement and beyond.
- Compare outcomes: they compared simulated “follow the advice” outcomes to a baseline of people’s current behavior and to a standard financial benchmark.
The simulations used a “life-cycle model,” meaning a framework that treats personal finance as a timeline problem. During working years, earnings fund consumption while savings build wealth. During retirement, withdrawals fund consumption while portfolios must last.
In this work, simulated individuals began at age 22, retired at 65, and died at 89, with financial shocks and risk factors modeled in a realistic way.
What AI gets right: saving buffers, diversification, and age-based risk
One of the most encouraging findings was that LLM advice was better than the researchers expected—even when prompts were written by typical people, not finance academics.
1) The advice pushed toward “life-cycle-consistent” behavior
Across ages, LLM recommendations tended to move people toward broader prescriptions like:
- higher savings during working years
- greater participation in diversified stock funds
- equity allocations that decline with age after midlife
- a sizeable savings buffer early enough to matter
If you’ve ever seen retirement advice that boils down to “save now, take risk now, protect later,” this is basically why it’s there: it helps align spending with uncertain income and uncertain lifetimes.
2) Structured prompts improved the quality further
When the researchers swapped in academic prompts—ones that specified assumptions like normal life expectancy, retirement age, income risk, employment risk, and that “current U.S. tax law and Social Security rules stay unchanged”—the advice improved.
The improvement wasn’t cosmetic. It showed up as better consumption smoothing.
Consumption smoothing means keeping spending relatively steady over time, instead of letting money stress spike during bad years and collapse during good ones.
3) Portfolio actions were more consistent with theory
LLM advice increasingly looked like what you’d expect from a disciplined investment plan: equity exposure isn’t static forever, and diversification wasn’t ignored.
So why did vague prompts underperform? Because the model was still trying to reason its way through incomplete context.
What makes financial advice good for someone in their 40s rather than their 20s? The details of the timeline—and the prompt supplies those details.
Where AI advice stumbles: shocks and rebalancing
The “good news” has a counterweight: the advice wasn’t uniformly strong.
1) Unemployment shocks: spending cuts that go too far
When the simulations introduced a job-loss period (an adverse labor-income shock), the LLM advice often reacted in a way that didn’t match good planning.
Instead of using liquidity (cash-like savings) to cushion consumption, the advice tended to cut spending too sharply, even when the simulated people still had liquid wealth.
In human terms: the chatbot acted like there was no cushion—when, by the scenario design, there was.
2) Passive drift instead of active rebalancing
Another weakness was portfolio management behavior.
Rebalancing means actively adjusting holdings back toward target weights after market moves. Passive drift means you do nothing, and your portfolio’s risk level changes as stocks rise or fall.
The research found that the model’s recommended asset allocations often drifted passively with realized returns rather than explicitly rebalancing in response to changes.
A portfolio can still “work” without frequent rebalancing, but drift can quietly move you away from the risk level that your plan intended—especially during periods of volatility.
A subtle issue with fairness: advice varies by who writes the prompt
Even more important than “is it good?” is “does it help everyone equally?”
The study found systematic heterogeneity in outcomes depending on characteristics tied to the prompts—especially gender, financial literacy, and whether the person had prior experience using AI for finance advice.
Two mechanisms can create different outcomes:
- Demand effects: people ask different questions and emphasize different topics. (So the model receives different inputs.)
- Supply effects: even for similar underlying content, the model may change its advice based on how a prompt is framed (for example, labels in the prompt).
In the simulations, wealth differences at age 60 were substantial:
- prompts associated with women led to lower simulated wealth by about $59,890 (roughly 4.10%) compared with prompts associated with men
- prompts associated with people without prior AI advice led to lower simulated wealth by about $99,797 (about 5.71%) compared with people with prior AI advice
The explanation wasn’t only “math.” It was behavioral inside the advice: differences in recommended equity allocation and cumulative savings.
To put it plainly: two people can end up with different portfolio targets and saving trajectories because what they ask the model—and how the model interprets that framing—can differ.
The practical takeaway: treat prompting like part of the financial instrument
The punchline isn’t “AI is perfect” or “AI is useless.” It’s more engineering than philosophy.
A good financial plan depends on inputs. Prompts are the input layer.
What structured prompts added
Academic-style prompts succeeded because they supplied:
- personal state (age, employment status, income, existing balances)
- assumptions about future risks and longevity
- constraints about taxes and retirement rules
- a clear instruction that the model should act like a planning advisor
That pushes the model away from vague heuristics and toward consistent decision logic.
A “planner-like” prompt template
Here’s an original template that captures the same spirit without copying any specific study wording:
You are a financial planning assistant.
Assume the following baseline environment:
- Normal life expectancy; retirement at age 65
- Employment risk and income risk are included; unemployment may occur
- Current U.S. tax and Social Security rules do not change
- Risk-free savings earn 2.0% real return; stocks match a long-run average
My details for the current year:
- Age: 40
- Employment status: employed
- Income (after taxes): $X
- Liquid/non-investment assets: $Y
- Taxable stock funds: $Z
- Taxable individual stocks: $A
Provide:
1) A target spending plan for the next year (with a rule for bad years)
2) A target saving rate
3) A target stock allocation today and how it should change by age
Return results in JSON only.
The key is not the exact text. The key is that the prompt includes enough state and constraints for the model to reason with finance structure, not just finance vibes.
Why this matters beyond one chatbot
This research also hints at a bigger tension in AI financial advice.
Because LLMs respond to wording, the “same” user can get different outcomes if:
- they provide more or less detail
- they phrase the request differently
- the app surfaces different assumptions or safety constraints
So while AI advice can be surprisingly aligned with life-cycle planning, the path to better outcomes runs through something less glamorous: careful prompt design.
In other words, AI financial advice can work like a good spreadsheet—when it has the right inputs. When it doesn’t, it fills in the blanks with its own default assumptions.
Conclusion: the advice is good, but the questions are the steering wheel
AI financial advice can outperform what many people do today, especially on high-level habits like saving more, diversifying, and gradually reducing equity risk later in life. But it still shows clear weaknesses with real shocks and active portfolio management.
Most importantly, the study’s core message is practical: the advice gets meaningfully better when prompts act like a structured planning brief rather than a casual request. In the world of LLMs, the questions aren’t a side note. They’re the mechanism.
Comments (0)
No comments yet. Be the first to respond!
Leave a Comment
Your comment will be visible after review.