Covariate-Adjusted Residual Policy Learning (CAR-PL) is introduced, an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support and support objective-specific ranking of SMB financial guidance from multi-action accounting logs.
Abstract
Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.
The results support the feasibility of direct language-model token generation for financial numerical prediction and decision-making, while motivating broader tests across assets, regimes, and random seeds.
It is demonstrated that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.
Ofir Ben Shoham, Shrutendra Harsola, Vignesh T. Subrahmaniam et al.· 0 citations
The end-to-end treatment policy delivered a statistically significant $+7.20\% lift in the primary long-term-value metric, demonstrating the feasibility of production-scale causal optimization under business constraints.
Changshuai Wei, John Bencina, Phuc Nguyen et al.· 0 citations
Out-of-time evidence indicates that tree-ensemble methods provide the strongest combination of ranking performance and probability accuracy, with the random forest providing the best out-of-time performance among the evaluated models, with reasonable discrimination and the lowest probability error.
Tuyen Le Nam, Tam Phan Huy· Journal of International Com...· 0 citations
AGT is best on all 13 KPIs against LightGBM, TimeMixer, and SOFTS in the matched seed-42 comparison, while final-architecture ablations show that relational attention, accounting topology, and the recency path each improve validation and test accuracy.
Shrutendra Harsola, Vignesh T. Subrahmaniam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.