Sep 2026· Journal of Advanced Computational Intelligence and Intelligent Informatics· Vol 30, pp. 1534-1554· 0 citations· 15 references
TL;DR
A unified information-theoretic framework for dimensionality reduction of finite distributions is provided, clarifying the structural role of f -divergence-based entropies and negentropies and their relation to classical information measures.
Abstract
Distributions are ubiquitous across scientific disciplines, extending well beyond probability and statistics. In machine learning, finite probability distributions arise naturally as the softmax output layers of convolutional neural networks and large language models, where they encode class probabilities in image classification and token probabilities in language generation. When such distributions are high dimensional, however, storage and computational costs become significant. In previous work, we introduced two families of generalized entropies derived from
f
-divergences, using majorization as a reference framework for comparing distributional homogeneity. In this paper, we extend that framework in several directions. First, we study majorization relationships between subcompositions of a distribution. Second, we introduce generalized negentropies derived from
f
-divergences and analyze their role alongside entropies in dimensionality reduction. Third, we embed both entropies and negentropies into the setting of Shannon’s information channel and show that the classical channel identities are satisfied exclusively by Shannon entropy. These results provide a unified information-theoretic framework for dimensionality reduction of finite distributions, clarifying the structural role of
f
-divergence-based entropies and negentropies and their relation to classical information measures.
After a historical introduction, the most important classical and quantum entropies are introduced as constructions in classical and quantum probability theory. Classical entropies are studied from large deviation theory, including theorems of Sanov, Cram\'er, G\"artner-Ellis, and Varadhan, and are illustrated in some...
TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components is proposed, indicating that activation-level degree control transfers to more general vision workloads.
Daehwa Ko, Jae-Hwan Kim, Seunghyun Ham et al.· 0 citations
Entropy is one of the most fundamental concepts in physics and information theory (Claude E. Shannon (1948), \cite{Shannon}), its correct understanding is essential for the study of various physical systems, especially in thermodynamics, statistical mechanics, and in thermal process engineering. This blending of conc...
Sergio Sánchez-Sánchez· European journal of physics· 0 citations
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collaps...
This work establishes a minimax lower bound under the score-entropy loss, and proposes an MLE-based thresholding estimator that matches this lower bound up to constant and polylogarithmic factors that depend on neighboring density ratios.
Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single soften...
G. Verbii, Ju-Ho Lee· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.