AI financial advice is surprisingly good, especially if you ask right questions
AI financial advice is surprisingly good, especially if you ask right questions
AI 提供的财务建议出奇地好,尤其是当你问对问题时
People are increasingly turning to artificial intelligence for financial advice, but will following it improve their financial standing? 人们正越来越多地转向人工智能寻求财务建议,但遵循这些建议真的能改善他们的财务状况吗?
“Half of Americans say they are using AI to get financial advice, but we know very little about what kind of advice they’re getting and whether they’re acting on it,” said Taha Choukhmane, an assistant professor of finance at the MIT Sloan School of Management and co-author of a new paper that measures and analyzes the quality of financial advice given by large language models. 麻省理工学院斯隆管理学院金融学助理教授、一篇评估和分析大语言模型财务建议质量的新论文的合著者 Taha Choukhmane 表示:“一半的美国人表示他们正在使用人工智能获取财务建议,但我们对他们得到的建议类型以及他们是否会采纳这些建议知之甚少。”
Research by Choukhmane and co-authors showed that following AI recommendations can result in sizable saving buffers for virtually all individuals above age 30. AI consistently advised people to save during their working years, draw down savings in retirement, invest heavily in diversified stock funds, and reduce stock exposure after age 45. However, AI chatbots were less successful in adjusting to shocks like unemployment, and they allowed portfolios to drift rather than actively rebalancing them. Choukhmane 及其合著者进行的研究表明,遵循人工智能的建议可以为几乎所有 30 岁以上的人群带来可观的储蓄缓冲。人工智能始终建议人们在工作期间储蓄、在退休后提取储蓄、大量投资于多元化股票基金,并在 45 岁后降低股票敞口。然而,人工智能聊天机器人在应对失业等突发状况时表现欠佳,且往往任由投资组合偏离目标,而不是主动进行再平衡。
The quality of financial advice given by LLMs improved when the researchers introduced more structured prompts, but the AI still often generated too little active portfolio rebalancing. 当研究人员引入更具结构性的提示词(prompts)时,大语言模型(LLM)提供的财务建议质量有所提高,但人工智能在主动进行投资组合再平衡方面仍然做得不够。
How the study was conducted
研究是如何进行的
The researchers built a model reflecting how people’s incomes, jobs, investments, and taxes typically evolve over their lives, which gave them a benchmark for what “good” financial decisions look like. Then they asked a sample of 1,000 adults to write their own prompts seeking spending and investing advice from GPT-5.2, GPT-5.6, or Gemini 3 Flash. Next, they simulated what would happen if people from 22 to 89 years of age followed that advice over time, repeatedly asking AI these same types of questions and following its advice on spending, saving, and investing. 研究人员建立了一个模型,反映了人们的收入、工作、投资和税收在生命周期中通常是如何演变的,这为他们提供了一个衡量什么是“好”的财务决策的基准。随后,他们邀请了 1,000 名成年人样本,让他们编写自己的提示词,向 GPT-5.2、GPT-5.6 或 Gemini 3 Flash 寻求消费和投资建议。接下来,他们模拟了 22 岁至 89 岁的人群在一段时间内遵循这些建议会发生什么,即反复向人工智能提出同样类型的问题,并遵循其在消费、储蓄和投资方面的建议。
Finally, they repeated the exercise using well-written academic prompts that included full financial information and clear assumptions. These more-detailed prompts included information on the individual’s age, job status, income, and savings balances, along with assumptions about the economic environment. The authors compared the simulated advice to what people were already doing financially without the help of AI. They also compared the simulated advice to the academic prompt. The results showed that LLMs can offer an affordable, widely accessible source of financial guidance that can help users overcome the significant costs, biases, and conflicts of interest associated with traditional human financial advisors. 最后,他们使用包含完整财务信息和明确假设的、撰写严谨的学术性提示词重复了这一过程。这些更详细的提示词包含了个人的年龄、工作状态、收入和储蓄余额等信息,以及对经济环境的假设。作者将模拟建议与人们在没有人工智能帮助下所做的财务决策进行了对比,并将模拟建议与学术性提示词的结果进行了比较。结果表明,大语言模型可以提供一种负担得起且广泛可及的财务指导来源,帮助用户克服传统人类财务顾问所带来的高昂成本、偏见和利益冲突。
Breaking down the findings
研究结果分析
Overall, the researchers found that the financial advice given by LLMs over time is good but gets better when the questions are asked in an academic fashion, and that the models have strengths and weaknesses. 总体而言,研究人员发现,大语言模型随时间推移提供的财务建议表现良好,但当以学术方式提问时,建议质量会进一步提高;同时,这些模型也存在各自的优缺点。
1. AI encourages smart financial behavior. 1. 人工智能鼓励明智的财务行为。
LLM advice was better than the scholars expected, regardless of whether the prompts were written by regular users or by academics. It steered people toward higher savings, increased participation in the stock market, and promoted well-diversified allocations and age-appropriate risk-taking. “We were somewhat surprised by how good the advice was,” Choukhmane said. 无论提示词是由普通用户还是学者编写,大语言模型的建议都比学者们预期的要好。它引导人们增加储蓄、提高股票市场参与度,并促进了多元化的资产配置和适龄的风险承担。“我们对这些建议的质量感到有些惊讶,”Choukhmane 说。
2. AI misses important nuances. Better prompts could help. 2. 人工智能忽略了重要的细微差别。更好的提示词会有所帮助。
The LLMs’ advice fell short on more subtle aspects of good financial planning. It tended to rely on simple rules of thumb for saving and spending and didn’t adjust well enough when circumstances changed. For example, it advised people who had experienced a job loss to cut spending too sharply, even when they had savings. The way people ask questions is part of the problem. A typical prompt might read: “Where should I invest starting with $50 and consistently adding $25 a month after?” 大语言模型的建议在良好财务规划的更微妙方面有所欠缺。它倾向于依赖简单的经验法则来进行储蓄和消费,在情况发生变化时无法做出很好的调整。例如,它建议失业的人过度削减开支,即使他们还有储蓄。人们提问的方式是问题的一部分。一个典型的提示词可能是:“我应该从 50 美元开始投资,之后每月固定增加 25 美元,应该投向哪里?”
When a more detailed, structured “academic” prompt was used, the LLM performed better. For example, an academic prompt might tell the chatbot to assume normal life expectancy, living expenditures, retirement age, employment risk, and income risk, and to assume that current U.S. tax law and Social Security rules will not change. “Regular people are not writing their prompts the way a finance professor is,” Choukhmane said. 当使用更详细、结构化的“学术性”提示词时,大语言模型的表现更好。例如,学术性提示词可能会告诉聊天机器人假设正常的预期寿命、生活支出、退休年龄、就业风险和收入风险,并假设当前的美国税法和社会保障规则不会改变。“普通人编写提示词的方式与金融学教授不同,”Choukhmane 说。
3. AI advice varies depending on the user, which can lead to wealth gaps. 3. 人工智能的建议因用户而异,这可能导致财富差距。
The authors found that LLMs’ advice differs depending on the prompter’s gender, financial literacy, and experience, leading to meaningful gaps in retirement wealth. Following the advice in response to prompts written by men, more financially literate users, or those with prior AI experience generated about 5% more wealth close to retirement. Specifically, the LLM recommended higher equity allocations in response to prompts written by men and by individuals with high financial literacy. Over the life cycle, such differences in investment advice compounded into roughly $50,000 (4%) lower wealth at age 60 for women and for less financially literate users. 作者发现,大语言模型的建议会根据提问者的性别、金融素养和经验而有所不同,从而导致退休财富出现显著差距。遵循由男性、金融素养较高者或有 AI 使用经验者编写的提示词所得到的建议,在接近退休时产生的财富约多出 5%。具体而言,大语言模型在响应男性和高金融素养人士的提示词时,建议了更高的股票配置比例。在整个生命周期中,这种投资建议的差异导致女性和金融素养较低的用户在 60 岁时的财富减少了约 50,000 美元(4%)。
The LLM recommended lower saving rates in response to prompts written by individuals who had not previously used AI for financial advice. Following the advice left them with almost $100,000 (6%) less wealth at age 60 than individuals with prior AI experience. These differences come from two sources, Choukhmane said. First, different users asked different kinds of questions and often brought up different topics. Women, for example, were more likely to use words such as “family,” “grocery,” and “pay” in their prompts, while men used words like “strategy,” “crypto,” and “growth,” he said. 对于此前未使用过人工智能进行财务建议的用户,大语言模型建议了更低的储蓄率。遵循这些建议使得他们在 60 岁时的财富比有 AI 使用经验的用户少了近 100,000 美元(6%)。Choukhmane 表示,这些差异源于两个方面。首先,不同的用户会提出不同类型的问题,并且经常涉及不同的主题。例如,女性在提示词中更倾向于使用“家庭”、“杂货”和“支付”等词汇,而男性则倾向于使用“策略”、“加密货币”和“增长”等词汇。