Skip to content

Similar papers

Open access Aug 2026

Restricted Boltzmann Machines, Bernoulli Mixtures and Sum-Product Networks: A Matched-Capacity Comparison on Binarized Image Data

Generative probabilistic models differ in a fundamental way that is rarely measured directly: some permit exact inference, while others are more expressive but require their likelihood to be estimated. This study compares three model families on binarized Fashion-MNIST and MNIST under identical preprocessing, identical...

Monther Ahmad, Syed Ejaz Ahmed · 0 citations
#machine learning Preprint Sep 2026

Understanding and Exploiting Anisotropy in Post-Training

LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large a...

Samyak Jha, Harshvardhan Saini, Yi-Zhen Liao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

Whitening a foundation-model embedding and using its squared norm as a training-free likelihood surrogate is motivated by the observation that whitened coordinates often appear approximately standard normal. We show that this observation follows from the projection central limit theorem and therefore does not imply a G...

Mohammed Ahnouch, Lotfi Elaachak · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation

This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigmoid relaxation makes...

Mariano Rivera · 0 citations
#machine learning Preprint Sep 2026

Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and...

Ning-Kang Peng, Qian-Feng Yu, Jing Mao et al. · 0 citations
#machine learning Preprint Sep 2026

Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation

For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this f...

Levin David Schwab · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.