Compositional Consistency as a Diagnostic Axis: How Operadic Structure, Analogical Retrieval, Latent Reasoning, Constructional Semantics, and Neural Representational Geometry Jointly Constrain Multi-Step Inference in Language Models
Abstract
This version corrects a citation error. Five passages on chaotic regularization in recurrent spiking networks cited the identifier 2606.04428, which belongs to an unrelated astrophysics paper on fast-spinning black holes; the passages describe arXiv:2606.04426 (Discrete signaling mediates chaotic regularization in recurrent neural networks), and all five now cite it. The wording of the claims is unchanged. Mathematical symbols that the previous PDF failed to render are now typeset. This version corrects specific citation errors found by an automated check and confirmed by hand; it has not had a full claim-by-claim audit. The full list of corrections is at the top of the PDF. Multi-step inference in large language models (LLMs) is typically evaluated at the level of final-answer accuracy, a metric that conflates many distinct failure modes: failures of decomposition, failures of integration, failures of analogical transfer, and failures of representational fidelity. This paper proposes a candidate framework, offered as a heuristic reading, not a formal derivation, in which compositional consistency serves as a unifying diagnostic axis for understanding where and why multi-step inference breaks down. We synthesize five to seven findings from recent arXiv preprints spanning cs.CL and q-bio.NC: (1) operadic consistency as a label-free, decomposition-sensitive accuracy predictor across twelve LLMs (arXiv:2606.13649, arXiv:2606.13634); (2) reasoning-aware analogical retrieval as a complementary training signal orthogonal to reward design (arXiv:2606.13680); (3) latent continuous-thought frameworks that aim to preserve probabilistic structure while bypassing discrete verbalization bottlenecks (arXiv:2606.06447); (4) constructional semantics acquisition as evidence that form-meaning compositionality emerges late and correlates with broader world-knowledge gains (arXiv:2605.31586); (5) neural representational geometry findings showing that chaotic recurrent networks produce smooth global population codes despite local roughness (arXiv:2606.04426); and (6) the variance allocation problem in brain foundation models, where higher-order co-skewness statistics, not captured by standard pretraining objectives, predict cognition better than second-order covariance alone (arXiv:2606.04010). Together these findings suggest, heuristically, that compositional consistency may track a structural property of inference simultaneously relevant to neural population geometry, language acquisition dynamics, and retrieval-augmented reasoning, though the cross-domain bridges are analogical, not mechanistic. The primary falsification path: if operadic consistency scores fail to predict accuracy on tasks requiring non-decomposable holistic inference (e.g., gestalt perceptual reasoning), the proposed axis would require revision or restriction to decomposition-structured tasks. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-13, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. Cited arXiv preprints: arXiv:2605.31586, arXiv:2606.04010, arXiv:2606.04426, arXiv:2606.06447, arXiv:2606.09767, arXiv:2606.11555, arXiv:2606.12342, arXiv:2606.13017, arXiv:2606.13634, arXiv:2606.13647, arXiv:2606.13649, arXiv:2606.13680 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.