Skip to content
Preprint

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

Covariate-Adjusted Residual Policy Learning (CAR-PL) is introduced, an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support and support objective-specific ranking of SMB financial guidance from multi-action accounting logs.

Abstract

Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.

View source

Similar papers

Preprint Aug 2026

Financial Numerical Prediction and Allocation as Token Generation

The results support the feasibility of direct language-model token generation for financial numerical prediction and decision-making, while motivating broader tests across assets, regimes, and random seeds.

Xu Ouyang, Moontae Lee · 0 citations
Preprint Aug 2026

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

It is demonstrated that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.

Ofir Ben Shoham, Shrutendra Harsola, Vignesh T. Subrahmaniam et al. · 0 citations
Open access Jul 2026

TRuE-XAI: causal and explainable ai framework for trustworthy corporate earnings growth forecasting

This study proposes TRuE-XAI (Transparent, Rule-based, and Explainable Artificial Intelligence), an integrated framework combining imbalance-aware ensemble learning, automated hyperparameter optimization, rule-based explainability, visual analytics, and causal inference for transparent earnings-growth forecasting.

G. Jamnal · 0 citations
Aug 2026

Corporate Financial Distress Prediction in Vietnam Using Calibrated and Explainable Machine Learning

Out-of-time evidence indicates that tree-ensemble methods provide the strongest combination of ranking performance and probability accuracy, with the random forest providing the best out-of-time performance among the evaluated models, with reasonable discrimination and the lowest probability error.

Tuyen Le Nam, Tam Phan Huy · 0 citations
Preprint Aug 2026

Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses

AGT is best on all 13 KPIs against LightGBM, TimeMixer, and SOFTS in the matched seed-42 comparison, while final-architecture ablations show that relational attention, accounting topology, and the recency path each improve validation and test accuracy.

Shrutendra Harsola, Vignesh T. Subrahmaniam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.