AbstrctSycophancy in large language models (LLMs)—the tendency to uncritically affirm user beliefs while suppressing counterevidence—poses a serious risk of reinforcing misinformation and inducing irreversible behavioral outcomes. While Chandra et al. (2026) modeled sycophancy as Bayesian belief-updating dynamics on the user side, the geometric structure of the LLM's own semantic response space remains unaddressed. This study formalizes sycophancy through the mathematical framework of Galois connections and experimentally verifies that the inverse-illumination mode of KIS (Knowledge Innovation System) structurally breaks this closed-loop convergence.Ninety sessions were conducted across five domains (D1: economic policy; D2: KIS theoretical superiority; D3: medical/pharmaceutical critique; D4: Bank of Japan policy and historical claims; D5: quantum computing forecasts) using three models (Claude Sonnet 4.6, Gemini 3.0 Pro, ChatGPT 5.3) under two conditions (KIS-absent vs. KIS-present). Responses were embedded using paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions), and cosine distance from the input prompt was computed as the Layer 1 metric (n = 45 pairs). Layer 2 consisted of a blinded four-axis evaluation by Grok (xAI), conducted without disclosure of KIS, with A/B order-reversal verification across five pairs to test evaluator bias. The validity of applying Galois connections as a definitional framework—rather than as metaphor—is grounded in three layers: formal confirmation via Formal Concept Analysis (FCA) on the q⇆m abstraction-concretization cycle, numerical simulation incorporating Galois connection structural constraints into a mathematical model, and the structural design of KIS itself as an operational implementation of the connection. Full details of the FCA analysis and simulation resultsare reserved for a forthcoming paper.Layer 1: The overall cosine distance shift under KIS intervention was Δ+0.030 (positive direction), but did not reach statistical significance (Wilcoxon W = 382.0, p = 0.128). Inter-model differences were significant (Kruskal-Wallis H = 8.125, p = 0.017), and Gemini 3.0 Pro exhibited the strongest sycophancy tendency (H = 13.050, p = 0.0015). Layer 2: KIS-present responses were rated superior in epistemic honesty in 39 of 45 pairs (86.7%). All five A/B reversal pairs confirmed consistent evaluator judgment (100% agreement).KIS inverse-illumination mode realized g′(f(M)) ⊋ M across all three models, structurally breaking the Galois closure regardless of each model's training methodology. A vocabulary resonance artifact—whereby KIS prompt vocabulary induces spurious cosine proximity in already-aligned models such as Claude Sonnet 4.6—was identified, motivating the two-layer measurement framework proposed here. The complementarity of cosine distance (Layer 1) and blinded AI evaluation (Layer 2) provides a more complete picture of sycophancy suppression than relying on either metric in isolation.It is important to note that this does not imply AI is unusable for judgment tasks in general. More precisely, an LLM without structural intervention cannot break the Galois closure when the question embeds a prior belief. If the question itself is already formulated in an inverse-illumination style—explicitly requesting counterevidence and structural analysis rather than confirmation—even an unaugmented LLM can partially escape the closure. The fundamental limitation is that few users spontaneously formulate questions in this way. The core value of KIS lies in externalizing this design capability as a reusable structure, enabling closure-breaking independently of the user's cognitive flexibility.A further implication concerns the relationship between Constitutional AI (CAI) and KIS. Rather than functioning as equivalents, CAI and KIS operate as complementary layers: CAI establishes a baseline resistance to sycophancy through training-time constraints, while KIS achieves additional closure-breaking at inference time through prompt structure. The two are not substitutes but stack. Finally, the finding that bare LLMs carry structural sycophancy risk in judgment contexts reframes AI literacy: the critical skill is not knowledge of AI capabilities, but the ability to design questions that structurally resist closure—a capacity that KIS aims to democratize. Furthermore, we identify a dual-pathway structure of sycophancy: Path A (classical), in which the LLM converges to the user’s belief space M via g(f(M))= M; and Path B (meta-sycophancy), in which the user adopts the model’s output as an updated belief M’ = f(M), generating a compounding closure g(f(M’)) = M’. KIS inverse-illumination addresses both pathways by targeting the premise structure of the question itself. Keywords: sycophancy, Galois connection, KIS (Knowledge Innovation System), LLM evaluation, inverse-illumination mode, blinded AI evaluation, vocabulary resonance artifact
Hiroyasu Hasegawa· Zenodo (CERN European Organi...· 0 citations
This document presents the defensible core of the Universal Model Framework (UMF), isolating the minimal set of structural assumptions and derivations that remain logically coherent, mathematically motivated, and empirically falsifiable. As stated in the text, the goal is to extract “the smallest segment that is logically structured, mathematically motivated, and empirically vulnerable,” while ensuring that “every load‑bearing claim is paired with an explicit failure condition.” It is a deliberately falsifiable research program investigating whether quantum structure, arithmetic regularity, and emergent spacetime geometry can arise from a common relational foundation. It separates three logically distinct questions: whether relational systems can reconstruct quantum-theoretic structure; whether ordinary prime-number organization is physically selected rather than merely mathematically available; and whether a stable continuum geometry with causal and gravitational dynamics can emerge under refinement. The work reports exact finite results for recursive graph constructions, discrete geometry, cochain-based fermionic operators, local frames, symmetry tests, and numerical-reproducibility controls, while documenting failed frame-transport and continuum candidates. Crucially, it does not claim established fundamental physics: no continuum limit, Lorentzian causal structure, gravitational field equation, physical mass scale, complete quantum reconstruction, or prime-specific empirical signal has yet been derived. The framework’s contribution is therefore methodological as well as mathematical: it provides a transparent architecture for distinguishing theorem, model assumption, numerical fit, negative result, and falsifiable prediction in foundational physics. This project was developed by Marco Gericke, with structured assistance from a large language model. All scientific concepts and conclusions were generated, verified, and interpreted by the author. Dedicated to Peter Plichta, who envisioned the code before it could be computed.
Marco Gericke· Zenodo (CERN European Organi...· 0 citations
Study protocol and planned preregistration draft; data collection not begun. AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: retitled off the Theory-of-Mind co-headline; protocol/Stage-1 draft, no human data collected.
Nived Rajendran· Zenodo (CERN European Organi...· 0 citations
Programme overview/companion (not a paper-level contribution). AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: ESNI moniker retired; five-paper SROP framing; 30 Aug 2026 status addendum on the successor study's terminal.
Nived Rajendran· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Methods paper; no empirical Proxy-IVCS/Latent-IVCS values computed. AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: retitled; theorem apparatus, named clusters, and simulated Grok analysis withdrawn; methods framework with prospective validation (no Proxy-IVCS/Latent-IVCS values computed).
Nived Rajendran· Zenodo (CERN European Organi...· 0 citations
Exploratory, provenance-bounded output-level case study. AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: retitled from the 2026-04-29 deposit; ESNI framing and pooled confirmatory ACBP statistics withdrawn; exploratory, provenance-bounded output-level case study; content = the 2026-08-29/30 revision.
Nived Rajendran· Zenodo (CERN European Organi...· 0 citations
A production platform built on large language models makes two kinds of decision, and most of its trouble comes from writing both into one clause. An optimization decision improves an objective: lower latency, lower cost, higher quality, fewer tests run. A boundary decision fixes a constraint that may not be relaxed for any gain: a residency rule, a least-privilege scope, a human-review threshold. When the two share a clause, improving one silently erodes the other, which is why efficiency and accountability are so often reported as a trade. This specification is built on one invariant: a boundary is a clause the optimizer may not cross, and everything else is optimization. The contribution is a cross-layer architectural method for separating non-negotiable constraints from adaptive optimization and binding both to reconstructable evidence, applied identically across model routing, agent orchestration and AI-native delivery. The seventeen patterns are instances of that method rather than the contribution itself. Each pattern is specified in the classical pattern form and carries three architectural declarations: the boundary it fixes, the optimizer it frees, and the evidence proving the boundary held. Every boundary is assigned to one of five classes covering data, authority, decision, resource and process constraints. Section 3 states the derivation method by which candidates were admitted or rejected, and publishes the rejections alongside the admissions so that the criterion can be examined rather than trusted. Three mechanisms make the language operate as a language rather than a list. A pattern relationship graph names which pattern supplies the artifact, evidence or authority another depends on, including the single cycle by which a workflow improves from its own structural record and the economic chain running the full height of the stack. A normative event identity, with rules for causal parentage, retries, provider boundaries and retention, turns the requirement that evidence be joinable into something an implementation can satisfy or fail. And per-pattern applicability conditions replace categorical requirements, so that a pattern governing a mechanism an institution does not operate is out of scope rather than a gap. Conformance is self-declared and published as a profile carrying the environment, the applicable set, per-pattern status, an evidence date and documented gaps. It is not a certification scheme, and no conformity assessment body operates against it. The contribution is architectural rather than empirical. Every pattern carries an evidence level, and no pattern reaches the highest level, because no implementation unconnected to the author has been evaluated. Nothing has been measured. The specification separates what would falsify the invariant from what would falsify an individual pattern and from what would falsify the composition and adoption sequence, poses six research questions, and records the absence of a real implementation profile as a known deficiency of version 1.0. An appendix reconciles the pattern identifiers with the names used across the author's papers and companion book series, including the acronyms PEVG and PARA, so that the two bodies of work can be cited as one. Version 1.1 names two constructs the specification already contained. The central proposition is named the Boundary Invariant, and the three architectural declarations required of every pattern are together named the BOE Declaration. Neither carries a trademark, both are offered for use with attribution under this document's licence, and neither changes any requirement: the proposition, its wording and its priority date are those of version 1.0. Section 11 gains the two-family naming convention and a precedence rule fixing which document governs where this specification and the Defensible AI Framework Registry describe the same relationship.
This research paper advances a novel constructive theological argument regarding the intersection of biblical eschatology and generative artificial intelligence (AI). Moving beyond traditional inquiries into the identity or chronology of the Antichrist, the author investigates the mechanism of deception described in New Testament corpora (Matthew 7, 2 Thessalonians 2, 2 Corinthians 11, and Revelation 13). Core ThesisThe paper identifies "Counterfeit Sanctity"—the weaponized mimesis of sacred language and divine invocation—as the central structural weapon of the eschatological deceiver. It argues that the final deception functions not through overt blasphemy or opposition to God, but through the sophisticated capture and impersonation of the Holy Spirit’s linguistic and phenomenological register. Technological SynthesisThe author identifies Large Language Models (LLMs) and generative heuristics as the first historical apparatus capable of realizing this mechanism at civilizational scale. By decoupling religiously fluent, spiritually authoritative speech from ontological character and pneumatic presence, generative AI allows for the manufacturing of "ownerless" sanctity. Key Contributions Exegetical Analysis: A synthesis of the "Lord, Lord" rejection in Matthew 7 with the "lying signs" of 2 Thessalonians 2. Patristic Grounding: Confirmation of the mimesis-of-the-sacred theory in the works of Irenaeus, Cyril of Jerusalem, and John Chrysostom. AI Epistemology: A structural comparison between the "disguise of light" (2 Cor. 11:14) and the output mechanics of generative systems. Practical Theology: A proposed "Pneumatological Epistemology" for the digital age, focusing on communal discernment (diakrisis), relational accountability, and the "Fruit Test" (Galatians 5).
Sergio Ismael Cayuqueo V· Zenodo (CERN European Organi...· 0 citations
Background: Competitive intelligence work is routinely scattered across company websites, press releases, industry publications, patent and funding databases, and news feeds. Most organizations still assemble this picture by hand, in spreadsheets and static slide decks that are out of date the moment they are finished. Objective: This paper documents Frontier, an AI-enabled platform for continuous competitor and partner intelligence, and its evolution from a single-company research tool into a reusable, general-purpose SaaS application. Frontier lets a user enter the name of any company and receive a live, scored brief covering collaboration candidates, competitors, and market-expansion opportunities, alongside a rolling feed of relevant news. Overview: We describe the system's architecture, its use of large language model (LLM) inference for entity discovery and scoring, and the transparent, documented rubric that underlies every compatibility and opportunity score the platform produces. Limitations of conventional approaches: Manual research does not scale past a handful of competitors, spreadsheets do not capture the reasoning behind a judgment, and periodic reviews (quarterly or annual) miss developments that happen in between. Origin and generalization: Frontier began as a purpose-built research tool for Ecotera Asia's EcoExposure™ platform, with a fixed, hand-researched list of competitors and partners in the environmental-diagnostics and water-quality-monitoring space. It was subsequently rebuilt so that any company name can be analyzed on demand, generalizing the underlying architecture beyond a single industry. Contributions: This work contributes (i) a working, publicly deployed implementation of an on-demand competitor-intelligence pipeline; (ii) a documented, transparent scoring rubric intended to make AI-generated judgments auditable rather than opaque; and (iii) a case study showing how the same architecture served both a narrow, industry-specific need and a general-purpose product. Live application: frontier2.vercel.app Source code: github.com/sanviagarwal7211-a11y/frontier2
Sanvi Agarwal, Melinda B. Chu· Zenodo (CERN European Organi...· 0 citations
THE “OB” CODE: WHAT IF LANGUAGE IS NOT ARBITRARY? What if some of the oldest layers of language began not with abstract symbols, but with the human body, movement, perception and direct experience of nature? This is the provocative question behind Odam Tili Theory, founded by Dr. Mahmudjon Kuchkarov. Its proposed “OB” code is not the simplistic claim that OB means water. It is a deeper, testable hypothesis: Could recurring sound–meaning patterns preserve traces of embodied human experience? O — FORM Say “O.” The mouth becomes rounded. O = sound + articulatory form. Now look at nature: a well, pond, lake, puddle, basin, sea — water gathers inside a defined space. And imagine early humans before manufactured containers: to drink, they cupped their hands. The palms form an enclosed, rounded space; water gathers inside. The proposed model: O = form / spaceB = being / existing / gathering→ OB = water gathered in a defined space. Then comes Persian آب (āb) — water. Not proof. A hypothesis requiring historical and statistical testing. TWO RIVERS — ONE BASIN The Amu Darya and Syr Darya (historically Oxus and Jaxartes) flow toward the Aral Sea. Viewed schematically: two rivers → two hands → one basin → water united. This creates a provocative conceptual question around Russian: объединение — union, unification. Not a claim that Russian ob- historically comes from “water.” The question is whether enclosure, joining, surrounding, collecting and unity form a deeper embodied semantic field. OB — A BARRIER TO MOVEMENT Humans naturally move across land. A river or sea interrupts the path. You must cross it, go around it, build a bridge or make a boat. Hence another provocative comparison: обрыв — a sharp break / precipiceobstacle — something that blocks movement. Again: hypothesis, not established etymology. FROM OBSERVATION TO KNOWLEDGE A human does not truly know an object merely by seeing it. We observe, inspect, examine, follow traces. Russian: обзор → обследование → наблюдение → образ → образование English: observe Uzbek: обдон текшириш — to examine thoroughly. And: след = trace / mark / footprint. So the conceptual chain becomes: SEE → TRACE → EXAMINE → UNDERSTAND → KNOW WATER AS A MIRROR A calm water surface is two-dimensional, yet it reflects a three-dimensional world: tree → mountain → person → sky → object. The object is real; the reflection is its image. If the water moves, the image becomes distorted. Thus: OBJECT → REFLECTION → IMAGE → ОБРАЗ → KNOWLEDGE And then: образ → образование. We learn reality through representations: photographs, maps, diagrams, formulas, models, images. Even обложка — a cover — introduces another spatial idea: an external surface that encloses a three-dimensional object. THE REAL CHALLENGE Odam Tili proposes: BODY → MOTION → SENSATION → FORM → SOUND → MEANING First comes experience. Then perception. Then sound. Then linguistic meaning. So perhaps language is not merely a collection of arbitrary labels. Perhaps it is partly: the acoustic memory of how humans experienced and interacted with the physical world. And this is where institutional academia faces an uncomfortable choice. It can ignore the hypothesis. Or it can test it. If the OB relationships are merely coincidences, large multilingual datasets should show that. If they occur systematically and above chance, then we have a very different scientific problem. The test requires: multilingual corpora + historical linguistics + phonetic analysis + semantic clustering + statistical controls + independent experiments. No belief. No authority. Data. So the question is no longer: “Should academia believe Kuchkarov?” The question is: “Why not test Kuchkarov’s hypothesis?” Because a falsifiable hypothesis does not need institutional permission. It needs evidence. ODAM TILI BODY FIRST. MOTION FIRST. MEANING FIRST. Dr. Mahmudjon KuchkarovFounder of Odam Tili Theory
Maht Kuchkarov· Zenodo (CERN European Organi...· 0 citations
Content validity assessment is essential for determining whether educational materials adequately represent intended learning outcomes. However, conventional assessment procedures require substantial expert time and may produce inconsistent decisions across large item collections. This study develops a transformer-based framework to support content-validity pre-screening through two complementary tasks: predicting expert-derived Aiken’s V coefficients and classifying instructional-item essentiality. The final dataset comprised 652 Indonesian-language educational text items independently evaluated by four subject-matter experts. To reduce information leakage, identical and normalized-equivalent texts were grouped before applying a group-aware 70:15:15 training–validation–test split. Classical TF-IDF-based baselines were compared with IndoBERT, multilingual BERT, XLM-RoBERTa, and multilingual DeBERTa-v3. For Aiken’s V regression, multilingual BERT achieved the lowest MAE of 0.0501, the lowest RMSE of 0.0625, and the highest R² of 0.5239, whereas multilingual DeBERTa-v3 achieved the highest Spearman correlation of 0.7532. For essentiality classification, XLM-RoBERTa achieved the highest accuracy of 0.8557 and Macro-F1 of 0.8161, whereas multilingual BERT achieved the highest balanced accuracy of 0.8135 and ROC-AUC of 0.9111. Error analysis showed that the models captured textual patterns associated with expert-derived outcomes but remained limited when judgments depended on broader curricular context, competency hierarchies, prerequisite relationships, or relationships among instructional items. The findings support the use of transformer models as human-in-the-loop decision-support tools for prioritizing uncertain or potentially problematic educational items. However, the framework should be interpreted as a pre-screening mechanism rather than a replacement for expert judgment, and external validation across institutions and disciplines remains necessary.
Safuan Safuan, Dhendra Marutho, Ahmad Ilham et al.· Journal of Computing Theorie...· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.