Skip to content
Open access

Evidence-Calibrated Financial Language Models for Macro-Policy Stance Classification and Decision Cards

Jul 2026 · Journal of Artificial Intelligence General science (JAIGS) ISSN:3006-4023 · Vol 10, pp. 1-23 · 0 citations · 85 references

TL;DR

Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms, so evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.

Abstract

Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms. This study evaluates hawkish, dovish, and neutral stance classification on all 496 FinBen-FOMC excerpts using leakage-controlled five-fold stratified-group cross-validation. Seven systems span a class-prior baseline, an n-gram language model, word-and-character TF-IDF, frozen FinBERT transfer scores, frozen MiniLM sentence embeddings, and two fused models. Four-fold inner cross-validation fits temperature and isotonic calibrators without access to outer-test labels. Evaluation combines macro-F1, Matthews correlation coefficient, expected calibration error, Brier score, confusion analysis, clustered bootstrap intervals, selective risk, and contrastive lexical evidence. TFIDF-LR achieved the highest pooled point estimates for macro-F1 (0.493) and MCC (0.246), followed closely by EvidenceFusion at 0.489 and 0.233. None of the macro-F1 differences between TFIDF-LR and another learned model was significant after Holm correction. Temperature scaling sharply reduced overconfidence in the n-gram and embedding systems; TFIDF+MiniLM attained the lowest temperature-scaled ECE (0.014), while TFIDF-LR achieved a Brier score of 0.587. Isotonic calibration lowered probability loss further for several models but reduced directional recall. Removing TFIDF-LR’s selected evidence terms reduced predicted-class confidence by 0.170 and changed 69.4% of labels, whereas matched random removal changed confidence by −0.019. At a 0.70 acceptance threshold, TFIDF-LR covered 6.45% of records with 21.9% risk, including a high-confidence error driven by tightening vocabulary despite explicit negation. Evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.  

Read PDF

Similar papers

Preprint Aug 2026

Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty

This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements. The project is explicitly inspired by iSent, Ita\'u's Central Bank sentiment classifier, particularly its sentence-level division of official communication into hawkish, dovish, neutral, and out-of-context classes. The implementation extends that idea in three directions. First, an LLM identifies short hawkish and dovish expressions and assigns each a 0-to-1 intensity weight. Second, the document index combines sentence counts with document-specific average signal intensities, producing a bounded score from -1 to 1. Third, a separate full-document layer measures forward-guidance direction, guidance explicitness, uncertainty level, and change in uncertainty. The empirical sample is restricted to communications dated August 2016 or later and contains 80 statements and 1,498 classified sentences from August 31, 2016 through August 5, 2026. Across this sample, 33.3% of sentences are hawkish, 18.0% dovish, 42.1% neutral, and 6.5% out of context. The average document score is +0.107, while the most hawkish reading is +0.570 in August 2021. The latest statement, dated August 5, 2026, scores +0.232, with eight hawkish, two dovish, and nine neutral sentences. Its structural overlay is more nuanced: guidance is directionally ambiguous but partly explicit, while uncertainty is classified as central and higher than at the prior meeting. Tone and the guidance-direction score have a contemporaneous Pearson correlation of 0.719. These are descriptive outputs, not a validated forecast of Selic decisions or DI returns. The main contribution is therefore methodological: a transparent, incremental, auditable system that separates rhetorical tone from policy guidance and uncertainty.

Gabriel de Macedo Santos · 0 citations
Open access Jul 2026

LLM-BASED SENTIMENT ANALYSIS FOR FINANCIAL DISTRESS DETECTION: EVIDENCE FROM THE 2023 U.S. BANK FAILURES

This paper evaluates whether large language model (LLM)-based sentiment analysis can detect financial distress more accurately than traditional dictionary-based methods. Using the 2023 U.S. bank failures as a natural experiment, Silicon Valley Bank (SVB), Signature Bank, and First Republic Bank each failed during March–May 2023, we construct monthly sentiment indices for five banks using an LLM alongside VADER, TextBlob, and FinBERT under an identical weighting framework. The LLM index consistently declines ahead of and during the failure period for the three distressed institutions while remaining stable for the two control banks (Bank of America, JPMorgan Chase). VADER, TextBlob, and FinBERT fail to detect the distress, remaining strongly positive throughout. Cohen’s Kappa coefficients near zero (0.01– 0.22) confirm that the methods capture fundamentally different signals. The LLM index is constructed using a severity-weighted aggregation scheme incorporating source credibility, model confidence, and recency, normalised via a tanh transformation. These findings suggest that LLMs interpret financial context rather than merely counting sentiment-bearing words, offering a meaningful advance for early-warning and financial risk monitoring applications.

N. Hasanli · 0 citations
Preprint Aug 2026

Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals

Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study an uncertainty-aware construction that feeds model-predicted risk -- decomposed into aleatoric and epistemic components -- directly into the covariance matrix of portfolio allocators, rather than treating portfolio risk as fixed or adjusting only expected returns. We evaluate the pipeline on Russell 2000 equities under three stock-selection regimes: a pure-alpha trigger that isolates abnormal stock moves not explained by macro indicators, a pure-beta trigger that captures macro-indicator moves before the stock itself fires, and a beta trigger in which both channels agree. Across the full holding-period grid, the separated pure-alpha and pure-beta legs usually dominate the beta intersection on Sharpe and return. Two horizons are especially informative. At one day, pure beta can work under low and moderate transaction costs because it captures immediate lead-lag spillovers from liquid macro and sector indicators into exposed small-cap stocks, but this advantage disappears at 100 bps when turnover and microstructure noise dominate. At 40 days, pure beta works for a different reason: slower macro repricing overtakes the firm-specific pure-alpha channel. The strongest conservative row is pure beta with GPT-4o mini sentiment, a Student-t target, a 40-day holding period, and risk parity allocation, reaching Sharpe 2.33 at 100 bps. The results suggest that stock-selection regime and allocator choice matter at least as much as the sentiment model, and that separating firm-specific and macro-exposure triggers is more informative than requiring both to fire simultaneously.

Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini et al. · 0 citations
#small language model Preprint Aug 2026

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

This work examines whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives and discusses implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

Sahab Zandi, Noah Kostesku, Christophe Mues et al. · 0 citations
Conference Open access 2026

Behavioral biases and artificial intelligence in banking decision-making: Toward explainable hybrid systems for SME financing

SME credit files arrive incomplete, and the gaps leave room for anchoring, confirmation bias and loss aversion. We compare human, algorithmic and hybrid credit decisions using a benchmark credit dataset alongside a vignette experiment with credit analysts working in Morocco's Souss-Massa region. The modelling arm pairs L2-regularised logistic regression with gradient-boosted trees, adding stratified validation, calibration analysis, SHAP and LIME. In the human arm, matched cases vary the requested amount while everything else is held constant. Analysts were least stable on borderline files, and their decisions moved with the anchor. The boosted model held steadier but leaned harder on indicators that track how thick a file is. AI-first assistance improved consistency and deepened deference to the model; human-first assistance preserved contextual overrides; explanation-gating struck the best balance, though only where SHAP and LIME agreed. We assess distribution through demographic-parity difference, disparate-impact ratio, equal-opportunity difference and false-positive-rate difference. What the results support is a governed hybrid: weak explanations withheld, overrides auditable, human review genuinely available. A regional sample and benchmark data bound how far any of these travels.

Hassan Ennaqui, Mohamed El Bourki, Abdellah Bakrim et al. · 0 citations