Skip to content

Scaling Laws, Tabular Data and Actuarial Ratemaking Models

Sep 2026 · 1 citation
Computer Science Economics

TL;DR

It is suggested that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.

Abstract

Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.

View source

Similar papers

#machine learning Preprint Sep 2026

SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws

Theoretical scaling laws from linear regression to multiclass classification to multiclass classification are extended, while connecting to empirical scaling laws observed in large-scale neural networks.

Konstantinos Christopher Tsiolis, Denny Wu, Christos Thrampoulidis et al. · 0 citations
#machine learning Preprint Sep 2026

When do data mixtures improve scaling laws? Insights from high-dimensional regression

Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely...

Di-Yuan Wu, Le-Han Chen, Theodor Misiakiewicz et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

This study establishes a first baseline and scaling procedure for the development of future OpenEuroLLM models, and characterize the dependence of loss on model capacity and dataset size, evaluating recently proposed scaling forms that explicitly model their interaction.

Niccolò Ajroldi, Diana-Alexandra Onutu, Haider Al-Tahan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Hyperparameter Scaling Laws Across MoE Sparsity

This work shows that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone.

Chang-Xin Tian, Kun-Long Chen, Jia Liu et al. · 0 citations
#machine learning Preprint Aug 2026

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

A tractable teacher--student model where a stable latent linear RNN generates trajectories and a sketched linear recurrent student is trained by safeguarded full-batch WSD gradient descent on next-token prediction is studied.

Zi-Yan Chen, Zhong-Zhu Zhou, Pei-Lin Liu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.