Skip to content
Preprint

How People Evaluate AI-, Expert-, and Peer-Style Financial Advice

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

Findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues, which position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.

Abstract

As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.

View source

Similar papers

Preprint Jul 2026

AI advice suppresses people's willingness to say"I don't know", even when the advice is wrong and accuracy is incentivized

Knowing when to say"I don't know"is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants'willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer.

Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro · 2 citations
Jul 2026

How generative AI behaves in the newsvendor problem: a behavioral experimental study

This study investigates behavioral biases of generative artificial intelligence (AI) models, specifically GPT-4o and Claude-Haiku-4.5, in inventory management using the newsvendor problem. This study compares AI decision-making with human-subject experiments to assess whether large language models (LLMs) replicate human cognitive bias and to identify prompt-design strategies that improve alignment with optimal outcomes. Controlled newsvendor experiments were conducted with generative AI models, mirroring established human-subject laboratory protocols. Prompt framing was systematically varied across three modifications: removing explicit waste and missed-profit information, simplifying instruction format and providing explicit optimization formulas. Results were benchmarked against normative economic predictions and existing human behavioral findings. Generative AI exhibits human-like human biases including risk aversion, loss aversion and demand chasing, but exhibits a stronger demand-chasing tendency than human participants. It responds to hypothetical incentives and displays bounded rationality. Prompt design significantly influences decision quality, producing decisions closer to theoretical benchmarks. This study empirically tests generative AI behavioral biases within a structured operations management experiment. It introduces a replicable methodology, extends findings across two architecturally distinct LLMs from different developers, and demonstrates that deliberate prompt design meaningfully reduces AI decision bias. The study also contributes a conceptual distinction between functionally analogous behavioral patterns and intrinsic psychological dispositions in LLMs, offering a more precise interpretive framework for AI decision-making research in operational contexts.

Jing-Jie Su, Yan Lang, Kay-Yut Chen · 0 citations
Review Jul 2026

Impact of AI-Powered Personal Finance Applications on Financial Decision-Making and Investment Behaviour Among Young Investors

Abstract The rapid proliferation of artificial intelligence (AI)-powered personal finance applications, spanning robo-advisors, automated budgeting tools, AI-driven expense trackers, and algorithmic investment platforms, has reshaped how young investors access financial information, form judgments, and act on investment opportunities. This paper presents a structured literature review and conceptual analysis of how AI-powered personal finance applications influence two closely related outcomes: financial decision-making quality and investment behaviour among young investors, broadly defined as individuals between eighteen and thirty-five years of age. Drawing on peer-reviewed research spanning behavioural finance, fintech adoption, robo-advisory literature, and digital financial literacy, the paper synthesizes evidence on the mechanisms through which AI personalization and automation affect cognitive biases, risk perception, financial confidence, and portfolio choices. A conceptual framework is proposed linking AI application features to decision-making and investment outcomes through mediating constructs such as perceived trust, financial self-efficacy, and perceived ease of use. The paper outlines a proposed mixed-methods research design for empirically testing the framework, discusses ethical and regulatory considerations including algorithmic transparency and data privacy, and identifies gaps in the current literature. The review indicates that AI-powered personal finance applications are consistently associated with improved financial confidence and more frequent, though not necessarily more diversified, investment activity among young users, with effects moderated by financial literacy, trust in automation, and platform design quality. The paper concludes with implications for fintech providers, regulators, and financial educators, and recommends directions for future longitudinal research. Keywords: Artificial intelligence; personal finance applications; robo-advisory; financial decision-making; investment behaviour; young investors; fintech; financial literacy

Shashank Adagond, D. R. G KARGAL · 0 citations
#artificial intelligence Preprint Aug 2026

How AI Prompts Can Teach Us About the Structure of Human Behavior

Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, it is found that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust.

Matthew O. Jackson, Benjamin S. Manning, Yutong Xie et al. · 0 citations
Preprint Aug 2026

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. We formulate financial advice generation as a reinforcement learning problem and fine-tune an open-weight language model using Group Relative Policy Optimization (GRPO). Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention. Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average Treatment Effect (CATE) estimator. Under this observational off-policy audit, our trained LLM achieves approximately twice the estimated gross-profit lift of the strongest evaluated commercial baseline ($0.0228$ vs.\ $0.0104$), together with the lowest downside rate and the least negative tail risk of any policy evaluated. Notably, the two evaluations do not rank the baselines identically: the untrained base model places last on the judge rubric but second on the causal audit, indicating that the audit captures a signal the judge does not. Our results demonstrate that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.

Ofir Ben Shoham, Shrutendra Harsola, Vignesh T. Subrahmaniam et al. · 0 citations
Preprint Aug 2026

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models'own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.

Syeda Anshrah Gillani, Mirza Samad Ahmed Baig · 0 citations