Skip to content

Category

large language models

256 papers

#large language models Dataset Open access Aug 2026

Investigating the Use of Large Language Models for Generating Abuser Stories for Early Security Threat Identification

This repository contains the complete dataset, experimental inputs, raw outputs, statistical scripts, and validation artifacts for the study assessing the effectiveness of Large Language Models (LLMs) in generating abuser stories from user stories under a constrained-context baseline. The study evaluated three lightweight models (GPT-4o mini, Claude 3.5 Haiku, Gemini 1.5 Flash) using two prompting techniques (Zero-Shot and One-Shot) across 10 real-world user stories.

Anonymous · 0 citations
#large language models Open access Aug 2026

There Are Three Ways to Break, and Strength Does Not Decide Which ── Buckling Is Decided by Shape, Fracture by a Flaw, Yielding by the Crystal ── One Section of One Material Has a Different Limit Once Its Length Changes ── [Paper 262]

The sentence “the strength of this material is 250 MPa” does not settle when it breaks. This paper asks what does settle it──the answer is that there are three ways to break and a different thing decides each. Buckling is decided by shape, fracture by a flaw, and yielding by the crystal. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──Euler's buckling load, slenderness, Griffith's fracture condition and the yield stress are all standard. No mechanics of materials is built──what is used is three formulas and one comparison. No value fit for design is given──no safety factor, no initial imperfection and no residual stress is included. The numbers are for seeing which of the three limits is lowest, not values for design. The fence on the elastic constants is not treated──Paper 261 treats -1<nu<1/2. This paper is what happens after one is seated inside that fence, up to breaking. No theory of plasticity is built──yielding is treated as one stress value, with no hardening and no flow rule. No value of the surface energy is claimed──gamma=1.0 J/m^2 is a posited value, not a measurement on a particular material. It is used to see orders of magnitude. Fatigue and time dependence are not treated──only a single loading is examined. Griffith is not confused with another──the Griffiths appearing in Paper 173 is an algebraic geometer and a different person from the A. A. Griffith of this paper. Relation to earlier papers: Paper 261 wrote that what raises the fence on the elastic constants is the positive definiteness of the energy──this paper treats what follows, namely where a material seated inside that fence breaks. Paper 190 measured rare on a logarithmic scale──this paper likewise writes the effect of a flaw as a square root and in orders of magnitude. Paper 196 counted “pressure” as four different quantities──this paper counts that “strength” is not even one quantity. Paper 201 counted “complete” as four different claims──the same shape of roll call. What is added is computing that the buckling stress of one section moves from 164.3 to 18.3 MPa on changing only the length, putting the switching slenderness at the concrete value 88.8577, confirming that a flaw tells as sqrt1000=31.6228, and setting the three limits in one table and writing that the lowest is the actual limit. First, change only the length of one section. A square steel column of side 31.6 mm buckles at 164.3 MPa when it is 1 m long and at 18.3 MPa when it is 3 m──nothing about the material has been changed (Section 2). Second, this is the core of the paper. The slenderness at which the mode changes is lambda_c=pisqrtE/sigma_y=88.8577──slimmer than this and buckling comes first, stubbier and yielding does, so one material has its limit exchanged (Section 3). Third, the third limit is set by a flaw. By Griffith's condition a flaw of 1 mum gives 356.8 MPa and one of 1 mm gives 11.3 MPa (Section 4). Fourth, a flaw tells as a square root. A flaw 1000 times larger lowers the strength by a factor of 31.6228──which is sqrt1000 itself (Section 4). Fifth, the three do not compete. Of three upper bounds that hold at once, the lowest is the actual limit──and which is lowest is decided not by the material but by shape and flaw (Section 5). Sixth, so strength is not a property of the material. The number sigma_y=250 MPa is held, and a column 3 m long still breaks at 18.3 MPa──a factor of 13.66 apart (Section 5). the sentence “the strength of this material is 250 MPa” did not settle when it breaks. There are three ways to break and a different thing decides each──buckling by shape, fracture by a flaw, yielding by the crystal. The three do not compete, and the lowest is the actual limit. Triple the length of one section and the buckling stress falls to a ninth; admit an invisible flaw of 10 mum and the fracture stress drops below half the yield. For a column 3 m long the 250 MPa on the data sheet stands a factor of 13.66 above the real limit and never gets its turn. One thing separates them──writing down which limit is being counted. Write it down, and the occasions for changing the material separate from those for changing the shape and those for removing the flaw. Do not write it down, and one goes on looking for a stronger steel for a column that breaks at 18.3 MPa. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 「この材料の強度は 250 MPa である」という一文は、壊れる条件を決めていない。本稿が問うのは、では何が決めているのかである──答は、壊れ方が三つあり、それぞれ別のものが決めているである。座屈は形が、破壊は傷が、降伏は結晶が決める。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──オイラーの座屈荷重、細長比、グリフィスの破壊条件、降伏応力は、いずれも標準的である。材料力学を作らない──使うのは三つの公式と、一つの比較だけである。設計に使える値を与えない──安全率も、初期不整も、残留応力も入れていない。数値は三つの限界の大小を見るためのものであり、実際の設計値ではない。弾性定数の柵を扱わない──論文261 が -1<nu<1/2 を扱う。本稿は柵の中に座ったあと、壊れるまでの話である。塑性論を作らない──降伏を一つの応力値として扱い、硬化も流れ則も扱わない。表面エネルギーの値を主張しない──gamma=1.0 J/m^2 は置いた値であり、特定の材料の測定値ではない。桁を見るために使う。疲労と時間依存を扱わない──一回の載荷だけを見る。グリフィスは別人と混同しない──論文173 に現れる Griffiths は代数幾何学者であり、本稿の A. A. Griffith とは別人である。既刊との関係:論文261 は弾性定数の柵を立てているのがエネルギーの正定値性だと書いた──本稿はその先、柵の中に座った材料がどこで壊れるかを扱う。論文190 は「稀」を対数の目盛りで測った──本稿も、傷の効き方を平方根と桁で書く。論文196 は「圧力」が四つの別の量であることを数えた──本稿は「強度」が一つの量ですらないことを数える。論文201 は「完備」が四つの別の主張であることを数えた──同じ形の点呼である。加えたのは同じ断面で長さだけを変えて座屈応力が 164.3 から 18.3 MPa まで動くことを計算したこと、切り替わりの細長比を 88.8577 と具体的な数で出したこと、傷の効き方が sqrt1000=31.6228 であることを確かめたこと、三つの限界を一つの表に並べ、最低のものが実際の限界になると書いたことである。 第一に、同じ断面で長さだけを変える。一辺 31.6 mm の正方形鋼柱は、1 m なら 164.3 MPa、3 m なら 18.3 MPa で座屈する──材料は一切変えていない(第2節)。 第二に、これが本稿の芯である。切り替わる細長比は lambda_c=pisqrtE/sigma_y=88.8577 である──これより細長ければ座屈が先、太短ければ降伏が先で、同じ材料で限界が入れ替わる(第3節)。 第三に、三つ目の限界は傷が決める。グリフィスの式では、傷が 1 mum で 356.8 MPa、1 mm で 11.3 MPa になる(第4節)。 第四に、傷は平方根で効く。傷を 1000 倍にすると強度は 31.6228 分の 1──sqrt1000 そのものである(第4節)。 第五に、三つは競合しない。同時に成り立つ三つの上限のうち、最も低いものが実際の限界になる──どれが最低かは、材料ではなく形と傷が決めている(第5節)。 第六に、だから強度は材料の性質ではない。 sigma_y=250 MPa という数を持っていても、長さ 3 m の柱は 18.3 MPa で壊れる──その差は 13.66 倍である(第5節)。 「この材料の強度は 250 MPa である」という一文は、壊れる条件を決めていなかった。壊れ方が三つあり、それぞれ別のものが決めているからである──座屈は形、破壊は傷、降伏は結晶。三つは競合せず、最も低いものが実際の限界になる。同じ断面で長さを 3 倍にすれば座屈応力は 9 分の 1 になり、10 mum の見えない傷が入れば破壊応力は降伏の半分以下に落ちる。材料表の 250 MPa は、長さ 3 m の柱については実際の限界の 13.66 倍上にあり、一度も出番が来ない。分けるものは一つ──どの限界を数えているのかを書き出すこと。書き出せば、材料を替えるべき場面と、形を変えるべき場面と、傷を消すべき場面が分かれる。書き出さなければ、18.3 MPa で壊れる柱に、より強い鋼を探し続けることになる。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#large language models Open access Aug 2026

The Length That Divides Gravity from Surface Tension Is Nobody's Size ── Water Gives 2.7273 mm ── Mercury Has 6.6621 Times the Surface Tension and Only 0.7009 Times the Capillary Length ── [Paper 309]

A liquid carries one length, l_c=sqrt(gamma/rho g). This paper asks whose size that length is──the answer is nobody’s. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──the Laplace pressure, Jurin’s law, the Bond number and the Rayleigh--Plateau instability are all standard. We do not build fluid mechanics──all we use is one length and one square root. We do not treat the contact angle──we compute with costheta=1 (complete wetting). Real contact angles, hysteresis and roughness are not treated. We do not treat dynamics──we say whether a column breaks, never how fast. We do not build plant physiology──whether the cohesion--tension theory is right is out of scope. We give only the number showing capillarity does not suffice. We do not build lung physiology──neither surfactant composition nor real alveolar shape is treated. Only the Laplace estimate. We claim no accuracy for representative values──gamma=0.0728 N/m and rho=998 kg/m^3 are values for showing orders and move with temperature. Relation to earlier papers: Paper 258 showed that 4pi appears only when the source is a point, using capillaries as its example of a cylindrical source──this paper stays in the same capillary setting but asks about a length rather than a solid angle. 258 treats potential, this paper treats an interface, and they share not one quantity. Paper 291 showed one sea holding three lengths──the l_c here is not a fourth length but a length that is the boundary between two forces. Paper 306 showed that the number of base units is a promise──l_c is not a promise: it comes out of two material quantities. Paper 140 showed that symmetry fixes ratios and dynamics fixes the scale──l_c is on the scale side. What is added is computing l_c for five liquids, showing that surface tension and capillary length reverse their order for mercury, checking that the Bond number equals (L/l_c)^2, counting that capillarity reaches only 0.74 m against a tree’s height, and pointing out that the break-up threshold 2pi R does not contain gamma. First, water gives 2.7273 mm. Mercury 1.9116 mm, ethanol 1.6977 mm, liquid helium 0.2625 mm (Section 2). Second, this is the core of the paper. Mercury has 6.6621 times the surface tension of water and a capillary length only 0.7009 times as long──its density is 13.5611 times greater, and only the square root of the ratio survives (Section 2). Third, the length is a boundary. The Bond number Bo=(L/l_c)^2 is exactly 1 there (Section 3). Fourth, capillarity is not what raises sap. In a 20 mum vessel it lifts water only 0.74 m (Section 4). Fifth, in an alveolus it is gamma that must move. Bare water gives 1456 Pa, surfactant 500 Pa (Section 5). Sixth, the separator is whether gamma enters the formula. The break-up threshold 2pi R does not contain it (Section 6). A liquid carries one length, l_c=sqrt(gamma/rho g), and for water it is 2.7273 mm──mercury 1.9116 mm, glycerol 2.2643 mm, ethanol 1.6977 mm, liquid helium 0.2625 mm. Mercury has 6.6621 times the surface tension of water and a capillary length only 0.7009 times as long──because its density is 13.5611 times greater, and l_c takes only the square root of the ratio. A factor of 4950 in gamma becomes a factor of 10.4 in l_c. The length is a boundary──the Bond number Bo=(L/l_c)^2 is exactly 1 at L=l_c; below it surface tension wins, above it gravity does. So l_c is nobody’s size──it is whom an object’s size is compared with. Capillarity is not what raises sap──in a 20 mum vessel it lifts 0.74 m, and 100 m would need a 0.1488 mum tube, 134 times narrower than any real one. In an alveolus what acts is not R but gamma──bare water gives 1456 Pa, surfactant 500 Pa. What removes the instability is not the value but the fact that gamma is not a constant. One thing separates them──whether gamma enters the formula. The break-up threshold 2pi R does not contain it and is pure geometry. Surface tension does not decide whether a column breaks, only how fast. *Revision Record Second edition (2026-08-30): The subject of this paper has been replaced. The first edition, titled “Living Tissue Where 4pi Appears, and Where It Does Not,” treated the solid angles of point, line and sheet sources and the diffusive reach around a capillary. That content duplicated Paper 258, “4pi Appears Only When the Source Is a Point”── even the numbers (24.1144 muV, 111.0508 muV, 316.78 mum, 44.80 mum) agreed, and Paper 258 has priority. The second edition stays in the same capillary setting but moves to a quantity 258 did not treat: the interfacial length l_c. All statements about solid angle have been removed from this paper; Paper 258 is the reference for them. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 液体には l_c=sqrt(gamma/rho g) という長さが一つある。本稿が問うのは、この長さは何の寸法かである──答は、どの物体の寸法でもないである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──ラプラス圧、ジュランの法則、ボンド数、レイリー=プラトー不安定は、いずれも標準的である。流体力学を作らない──使うのは一つの長さと、一つの平方根だけである。接触角を扱わない──costheta=1(完全濡れ)で計算する。実際の接触角、ヒステリシス、粗さの効果は扱わない。動的な現象を扱わない──切れるかどうかは書くが、どれだけ速く切れるかは書かない。植物生理を作らない──凝集力説の当否は扱わない。毛管だけでは足りないという数を出すだけである。肺の生理を作らない──界面活性剤の組成も、実際の肺胞の形も扱わない。ラプラス圧の見積りだけである。代表値に精度を主張しない──gamma=0.0728 N/m、rho=998 kg/m^3 は桁を見るための値であり、温度で動く。既刊との関係:論文258 は 4pi が出るのは源が点のときだけだと示し、円柱源の例として毛細血管を扱った──本稿は同じ毛細の場で、立体角ではなく長さを問う。258 が扱ったのは電位、本稿が扱うのは界面であり、共通の量を一つも持たない。論文291 は同じ海に三つの長さがあると示した──本稿の l_c は四つ目の長さではなく、二つの力の境目としての長さである。論文306 は基本単位の個数が約束だと示した──l_c は約束ではなく、材料の二つの量から出る。論文140 は対称性が比を決め力学が尺度を決めると示した──l_c は尺度の側である。加えたのは五つの液体で l_c を計算したこと、水銀で表面張力と毛管長の大小が逆転することを示したこと、ボンド数が (L/l_c)^2 に一致することを確かめたこと、木の高さに毛管が 0.74 m しか届かないと数えたこと、切れる閾値 2pi R に gamma が入らないと指摘したことである。 第一に、水では 2.7273 mm である。水銀 1.9116 mm、エタノール 1.6977 mm、液体ヘリウム 0.2625 mm(第2節)。 第二に、これが本稿の芯である。水銀は表面張力が水の 6.6621 倍なのに、毛管長は 0.7009 倍と短い──密度が 13.5611 倍だからであり、比の平方根しか効かない(第2節)。 第三に、この長さは境目である。ボンド数 Bo=(L/l_c)^2 が ちょうど 1 になる(第3節)。 第四に、木を登らせているのは毛管ではない。半径 20 mum の道管で、毛管が上げるのは 0.74 m だけである(第4節)。 第五に、肺胞では gamma を動かすほうが効く。裸の水なら 1456 Pa、界面活性剤で 500 Pa(第5節)。 第六に、分離子は「gamma が入るかどうか」である。液柱が切れる閾値 2pi R に gamma は入らない(第6節)。 液体には l_c=sqrt(gamma/rho g) という長さが一つあり、水では2.7273 mmである──水銀 1.9116 mm、グリセリン 2.2643 mm、エタノール 1.6977 mm、液体ヘリウム 0.2625 mm。水銀は表面張力が水の 6.6621 倍なのに、毛管長は 0.7009 倍と短い──密度が 13.5611 倍だからであり、l_c は比の平方根しか取らない。 gamma が 4950 倍ひらいても l_c は 10.4 倍しかひらかない。この長さは境目である──ボンド数 Bo=(L/l_c)^2 は L=l_c でちょうど 1 になり、下では表面張力が、上では重力が勝つ。だから l_c はどの物体の寸法でもない──物の大きさを比べる相手である。木を登らせているのは毛管ではない──半径 20 mum の道管で毛管が上げるのは0.74 mであり、100 m には0.1488 mum の管が要る(実際より 134 倍細い)。肺胞で効くのは R ではなく gamma である──裸の水なら 1456 Pa、界面活性剤で 500 Pa。不安定を消しているのは値ではなく、gamma が定数でないことである。分けるものは一つ──gamma が式に入るかどうか。液柱が切れる閾値 2pi R に gamma は入らず、純粋に幾何である。表面張力は「切れるかどうか」を決めず、「どれだけ速く切れるか」だけを決める。 *改訂記録 第2版(2026-08-30):本稿は主題を入れ替えた。 第1版は「4pi が出る生体と、出ない生体」と題し、点源・線源・面源の立体角と毛細血管の到達距離を扱っていた。 その内容は論文258「4pi が出るのは、源が点のときだけである」と重複していた── 数値(24.1144 muV・111.0508 muV・316.78 mum・44.80 mum)まで一致しており、 先行するのは論文258 である。第2版は同じ毛細の場に留まりつつ、 258 が扱わなかった量(界面の長さ l_c)に主題を移した。 立体角についての記述は本稿から全て削除し、論文258 を参照先とする。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#large language models Open access Aug 2026

Intermediate-Task Difficulty and Robustness in Zero-Shot Cross-Lingual Transfer on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does the selection of English intermediate-task difficulty influence the robustness of zero-shot cross-lingual transfer on XTREME-R when evaluated under adversarial perturbations (e.g., typos, paraphrasing) in target languages, measured by accuracy degradation rates? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.3/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Sycophancy as Galois Closure: How KIS Structurally Prevents Delusional Convergence in LLMs

AbstrctSycophancy in large language models (LLMs)—the tendency to uncritically affirm user beliefs while suppressing counterevidence—poses a serious risk of reinforcing misinformation and inducing irreversible behavioral outcomes. While Chandra et al. (2026) modeled sycophancy as Bayesian belief-updating dynamics on the user side, the geometric structure of the LLM's own semantic response space remains unaddressed. This study formalizes sycophancy through the mathematical framework of Galois connections and experimentally verifies that the inverse-illumination mode of KIS (Knowledge Innovation System) structurally breaks this closed-loop convergence.Ninety sessions were conducted across five domains (D1: economic policy; D2: KIS theoretical superiority; D3: medical/pharmaceutical critique; D4: Bank of Japan policy and historical claims; D5: quantum computing forecasts) using three models (Claude Sonnet 4.6, Gemini 3.0 Pro, ChatGPT 5.3) under two conditions (KIS-absent vs. KIS-present). Responses were embedded using paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions), and cosine distance from the input prompt was computed as the Layer 1 metric (n = 45 pairs). Layer 2 consisted of a blinded four-axis evaluation by Grok (xAI), conducted without disclosure of KIS, with A/B order-reversal verification across five pairs to test evaluator bias. The validity of applying Galois connections as a definitional framework—rather than as metaphor—is grounded in three layers: formal confirmation via Formal Concept Analysis (FCA) on the q⇆m abstraction-concretization cycle, numerical simulation incorporating Galois connection structural constraints into a mathematical model, and the structural design of KIS itself as an operational implementation of the connection. Full details of the FCA analysis and simulation resultsare reserved for a forthcoming paper.Layer 1: The overall cosine distance shift under KIS intervention was Δ+0.030 (positive direction), but did not reach statistical significance (Wilcoxon W = 382.0, p = 0.128). Inter-model differences were significant (Kruskal-Wallis H = 8.125, p = 0.017), and Gemini 3.0 Pro exhibited the strongest sycophancy tendency (H = 13.050, p = 0.0015). Layer 2: KIS-present responses were rated superior in epistemic honesty in 39 of 45 pairs (86.7%). All five A/B reversal pairs confirmed consistent evaluator judgment (100% agreement).KIS inverse-illumination mode realized g′(f(M)) ⊋ M across all three models, structurally breaking the Galois closure regardless of each model's training methodology. A vocabulary resonance artifact—whereby KIS prompt vocabulary induces spurious cosine proximity in already-aligned models such as Claude Sonnet 4.6—was identified, motivating the two-layer measurement framework proposed here. The complementarity of cosine distance (Layer 1) and blinded AI evaluation (Layer 2) provides a more complete picture of sycophancy suppression than relying on either metric in isolation.It is important to note that this does not imply AI is unusable for judgment tasks in general. More precisely, an LLM without structural intervention cannot break the Galois closure when the question embeds a prior belief. If the question itself is already formulated in an inverse-illumination style—explicitly requesting counterevidence and structural analysis rather than confirmation—even an unaugmented LLM can partially escape the closure. The fundamental limitation is that few users spontaneously formulate questions in this way. The core value of KIS lies in externalizing this design capability as a reusable structure, enabling closure-breaking independently of the user's cognitive flexibility.A further implication concerns the relationship between Constitutional AI (CAI) and KIS. Rather than functioning as equivalents, CAI and KIS operate as complementary layers: CAI establishes a baseline resistance to sycophancy through training-time constraints, while KIS achieves additional closure-breaking at inference time through prompt structure. The two are not substitutes but stack. Finally, the finding that bare LLMs carry structural sycophancy risk in judgment contexts reframes AI literacy: the critical skill is not knowledge of AI capabilities, but the ability to design questions that structurally resist closure—a capacity that KIS aims to democratize. Furthermore, we identify a dual-pathway structure of sycophancy: Path A (classical), in which the LLM converges to the user’s belief space M via g(f(M))= M; and Path B (meta-sycophancy), in which the user adopts the model’s output as an updated belief M’ = f(M), generating a compounding closure g(f(M’)) = M’. KIS inverse-illumination addresses both pathways by targeting the premise structure of the question itself. Keywords: sycophancy, Galois connection, KIS (Knowledge Innovation System), LLM evaluation, inverse-illumination mode, blinded AI evaluation, vocabulary resonance artifact

Hiroyasu Hasegawa · 0 citations
#large language models Open access Aug 2026

Oral MLLM Scoping Review Protocol: Multimodal Large Language Models in Stomatology

OSF 注册文案:口腔多模态大语言模型范围综述方案 1. 方案标题 多模态大语言模型在口腔医学中的应用:从影像诊断到智能病理——范围综述方案 Multimodal Large Language Models in Stomatology: From Imaging Diagnosis to Intelligent Pathology — A Scoping Review Protocol 2. 研究团队 刘雨、王翔、侯君、唐建军、翁雁鸣、马超群、董青山(通讯作者) 中国人民解放军中部战区总医院口腔科,湖北 武汉 3. 研究问题(PCC 框架) 人群(Population):接受口腔影像学或口腔病理学检查的患者、口腔临床诊疗场景,以及承担影像/病理判读的口腔专科医师(含初级与资深医师阅片场景)。 概念(Concept):多模态大语言模型(含专病视觉-语言模型与智能体式诊断系统)、牙科视觉基础模型,以及基于多实例学习(MIL)的全切片病理方法;涵盖模型构建、评测基准与临床验证三类研究。 情境(Context):口腔医学中的影像诊断(全景 X 线片、根尖片、头影测量片、CBCT、口内照片等 2D/3D 模态)与病理智能诊断,包括"研究验证"与"临床落地"两条路径。研究问题:①口腔专用 MLLM 及其方法学基础呈现何种技术路线与能力分布?②现有评测基准与临床验证证据的强度如何?③该领域存在哪些证据缺口与转化障碍? 4. 综述类型与报告规范 类型:范围综述(Scoping Review) 报告规范:PRISMA-ScR(Tricco 等,2018) 按设计不进行定量合并(Meta 分析);对效应量可比性与未来定量综合可行性作系统评估 5. 检索策略 数据库:arXiv、PubMed、Google Scholar、中国知网(CNKI) 时间窗:2025-01-01 至检索执行当日(计划 2026-08-30) 语言:中英文;限定标题与摘要(Google Scholar 以标题检索为主辅以人工过滤) PubMed 检索式:("oral"[Title/Abstract] OR "dental"[Title/Abstract] OR "dentistry"[Title/Abstract] OR "stomatology"[Title/Abstract]) AND ("multimodal"[Title/Abstract] OR "vision-language"[Title/Abstract] OR "large language model"[Title/Abstract] OR "MLLM"[Title/Abstract]) AND (diagnosis[Title/Abstract] OR imaging[Title/Abstract] OR pathology[Title/Abstract] OR panoramic[Title/Abstract] OR radiograph[Title/Abstract] OR benchmark[Title/Abstract]),时间限定 2025/01/01 至检索日 arXiv 检索式:all:(oral OR dental OR dentistry OR stomatology) AND all:(multimodal OR "vision-language") AND all:(model OR LLM OR foundation) 概念块:模型块(MLLM/VLM/LLM/foundation model)、领域块(oral/dental/dentistry/stomatology/口腔/牙科)、任务块(diagnosis/imaging/pathology/panoramic/radiograph/benchmark) 补充检索:评价工具(QUADAS-AI、TRIPOD+AI、CLAIM)、跨专科对照文献、口腔领域外奠基性方法学文献、WHO 官方报告,以及口腔领域内早于时间窗的奠基性背景文献与基准性资源(如 GBD、DENTEX)不受时间窗限制,经滚雪球检索补充;官方媒体与机构官网作为灰色文献仅记录产业动态。 6. 纳入与排除标准 纳入:①口腔/牙科专用 MLLM、牙科视觉基础模型或口腔评测基准;②直接相关的全切片病理 MIL 方法学工作;③口腔 MLLM 临床验证研究;④中英文文献。 排除:①单任务单模态 CNN 研究;②观点性文章、社论及无原始数据的方法学评论;③无可迁移方法学的非口腔文献;④重复发表。 7. 筛选与数据提取 两位作者(刘雨、王翔)独立筛选,初筛基于标题与摘要,复筛阅读全文,分歧协商解决 标准化表格提取:模型名称、年份、技术路线、基座/方法、数据规模与任务、关键验证结果、证据来源类别(同行评议/预印本/灰色文献);第三位作者(侯君)核对 全程记录各库命中数、去重数、初筛排除数、全文评估数、纳入数及排除原因 8. 证据分级与可比性评估 证据来源分类:同行评议/预印本/灰色文献,逐条标注(出版状态≠证据等级) 对唯一具备可提取验证设计的研究(DentVLM)按 QUADAS-AI 作非正式分域评价(author appraisal);报告完整性参照 TRIPOD+AI 与 CLAIM 核查 可比性框架:≥2 项同任务同指标且可提取效应量及 95% CI(或可重构 2×2 表)为可比性判定条件;≥3 项同质研究为执行随机效应合并的最低数量条件;漏斗图与 Egger 需 ≥10 项。无论条件是否满足,本综述均不执行合并 9. 预期产出 ①口腔 MLLM 领域证据地图;②效应量可提取性与可比性评估表;③临床验证"最小报告集"建议;④未来系统综述/Meta 分析的前提条件清单 10. 注册与备案声明 本方案于 2026-08-30 在 OSF 注册备案。检索执行、筛选与数据提取均在本注册之后进行,筛选计数将全程记录并纳入最终报告。

Liu Yu · 0 citations
#large language models Open access Aug 2026

Evaluating Technology Acceptance of Vernacular AI Interfaces: An Empirical Study Among Multilingual Engineering Students in Kasaragod

Generative Artificial Intelligence (AI) tools have become embedded in the everyday academic practice of undergraduate engineering students, yet most large language models remain optimised for standard English rather than the code-mixed, multilingual registers through which students in linguistically plural regions actually think and communicate. This study examines technology acceptance of vernacular and code-mixed AI interaction among 84 undergraduate engineering students enrolled in APJ Abdul Kalam Technological University (KTU)-affiliated institutions in Kasaragod district, Kerala, a region historically described as Saptha Bhasha Sangama Bhoomi, the confluence land of seven languages. Using a structured questionnaire grounded in the Technology Acceptance Model (Davis, 1989), the study measured Perceived Usefulness (PU), Perceived Ease of Use (PEOU), Output Accuracy, and Linguistic Inclusion across five research hypotheses. Findings indicate that students from regional-medium secondary schooling backgrounds report significantly higher vernacular or code-mixed AI prompting than English-medium peers, chi-square(3, N = 84) = 22.91, p < .001. Perceived Usefulness correlates strongly with Perceived Ease of Use, r = .64, p < .001. Students who habitually use vernacular or code-mixed prompts report significantly higher ease of use than strictly English prompters, t(82) = 2.01, p = .048. Perceived terminological distortion is positively associated with reported reliance on AI-translated academic content, r = .27, p = .012, and native speakers of the unscripted Tulu dialect report markedly higher AI comprehension failure than speakers of scripted regional languages, t(79) = 11.60, p < .001. The results support all five hypotheses and highlight a persistent linguistic-inclusion gap in generative AI systems used within multilingual engineering classrooms. Implications for dialect-aware AI design and inclusive digital pedagogy in polyglot regions such as Kasaragod are discussed.

Amal George · 0 citations
#large language models Open access Aug 2026

The Symmetric Unit and the Midline Theorem: A First-Principles Geometric Construction in which Classical ζ is the Limit of Finished Stations and the Riemann Hypothesis is the Statement that the Limit has all Non-Trivial Zeros on the Only Primary Self-Ratio Available

The Symmetric Unit and the Midline Theorem A first-principles geometric construction in which the only primary self-ratio on a finished segment is 1/2. Classical ζ is identified later as a derived unit of comparison on that cut. In this order the Riemann Hypothesis is the statement that the limit has every non-trivial zero there, because no other self-ratio is available without privilege. This deposit timestamps a construction that did not begin as an attempt on the Riemann Hypothesis. It began from a question about privilege in a live model: when everything is moving, what is allowed to count as “now”? The question was pursued through elementary geometry, inversion, and a refusal to appoint a second privileged count. The master sequence is the only count allowed to go ahead. Everything else is dated against a station that has already finished. Abstract. The construction produces a rigid stock of measurements along a master sequence that cannot skip ahead. The central object is the Symmetric Unit. Its permanent cut is the unique primary, scale-invariant, ±-equal self-description of location on a segment: the midpoint ratio 1/2. The same unit carries a rigid similar triangle and a local circle. Later stations inherit that package at a larger radius. Trace comes first. Measurement second. Unit language third. Place is a coincidence of readings. Location is a place after a unit has been asked. A station is a finished prime on the master sequence — a bridge, not a zero. A live self-measure such as nπ/2 is a magnitude owned by the walk in the walk’s own ratio. A zeta zero is a later question: a derived unit assembled from independent processes and compared from a named HERE. Those meetings, when they occur, still sit on the only primary cut both sides already hold. Classical ζ is named only after the stock is built. The Midline Theorem is the statement that the limit has every non-trivial zero on the pull-back of the midpoint ratio. Forced identities are proved. Named constructions are labeled. No completed prime list and no external π are imported. Numerics demonstrate rigidity of the stock at finite depth. They are not a search for zero locations. This record contains. The paper, v2 (PDF and TeX). The script that rebuilds the appendix drawings from the stock. The live SU model, which walks that stock, writes the same CSVs, and exports a 3D coil whose height is the 1/2 axis. An extended drawing set read from the walk. Numerics that show rigidity, not zeros. This record does not contain. A claim that a gap midpoint is a zero. A claim that a live arc — 9π/2 at the apex of [7, 11], or any other self-measure — is a zeta zero. Version 1 remains the closed timestamp of the first writing. This version is the same construction with the order of language tightened, the drawings rebuilt from the stock, and the live machine included so the rigidity can be inspected.

Justin Erholtz · 0 citations
#large language models Open access Aug 2026

memoria.ia: Resolutive Memory — v1.0.0 Release Candidate 1

Memoria.ia v1.0.0-rc1 — Release Candidate 1 Release date: 2026-08-30 Summary v1.0.0-rc1 is the first publication candidate for the Memoria.ia v1 line. It consolidates the validated Resolutive Memory research lineage with the deployable PC/server product layer and the native/mobile runtime path, while keeping post-v1 experimentation isolated from the release candidate. The release architecture remains: application / OFF.IA / agent ↓ Memoria.ia ↓ Resolutive-DB / BDR Memoria.ia owns memory semantics and state. Resolutive-DB owns durable persistence. Optional LLMs are consumers, not the authoritative memory store. Included capabilities persistent local-first memory state; organization and namespace isolation; provenance and authority lineage; conservative HIT / MISS / UNRESOLVED resolution; semantic, episodic, temporal and relation kernels; correction/supersession behavior with preserved lineage; PC/server FastAPI product boundary; Docker/Compose deployment; provider-neutral language-model adapters; metrics and context-selection instrumentation; integrity-checked backup/restore; native production runtime; Android arm64-v8a mobile ABI; durable native BDR persistence and restart recovery; indexed native resolution for large-memory workloads; reproducibility and release metadata gates; official Memoria.ia visual identity assets. Frozen candidate provenance The functional candidate was frozen at: dc73cbcdddfe20e0729e7e6bdea4697f7e8308cd That commit integrated PR #112, which preserved ranking, confidence, provenance policy, ABI and BDR contracts while adding the indexed native resolve lineage. The release branch adds publication metadata, version alignment, release documentation and current branding without importing post-v1 PR #116 runtime behavior. Validation evidence The exact functional lineage used for this release candidate passed the recorded required gates before release preparation: Android mobile ABI: PASS; native production image: PASS; Ubuntu/Windows candidate regression: PASS; BDR Linux/Ubuntu/Windows integration: PASS; native 100 / 1k / 10k benchmark matrix: PASS. Recorded 10k native resolve benchmark improvement versus the prior frozen baseline: p50: 693.233 ms -> 6.288 ms (~110x); p95: 710.630 ms -> 6.391 ms (~111x). These figures are environment- and workload-specific benchmark evidence, not universal latency guarantees. Publication metadata Release version: 1.0.0-rc1 Python package version: 1.0.0rc1 License: Resolutive Research and Non-Commercial License (RRNCL) v1.0 Author: Marcelo Roldão Matos ORCID: 0009-0003-6075-4680 RSMS compatibility: 1.0-rc.1 A new archival DOI should be assigned to this publication. The v0.95 DOI must not be reused as the release DOI for v1.0.0-rc1. Why this is RC1 rather than final v1.0 The repository currently declares compatibility with RSMS 1.0-rc.1, and the published Resolutive Science baseline remains on that release-candidate specification. Therefore Memoria.ia is published as v1.0.0-rc1 rather than claiming final v1.0 compatibility prematurely. Final v1.0 promotion requires: successful release-candidate metadata and regression gates; reproducibility from the public release state; compatibility re-audit against stable RSMS; no release-blocking regression found during RC use; final archival metadata and DOI synchronization. Explicitly excluded from RC1 The following post-v1 work is not part of this release candidate: external/public knowledge learning from OFF.IA Curiosity (issue #114 / PR #116); autonomous curiosity policy; new MA2A federation transport; multimodal post-v1 expansion; new semantic-consolidation phases from the post-v1 roadmap. Those features continue independently after this publication. Security boundary This release candidate is not represented as independently production-security certified. Authentication, isolation, integrity and negative-path controls exist and are tested, but no independent production security audit is claimed. Claims boundary This release does not claim: artificial general intelligence; biological equivalence; replacement of general-purpose LLMs; universal O(1) semantic resolution; production-ready MA2A federation; security certification. Claims are limited to the implementation, tests, benchmarks and reproducible evidence recorded in the repository.

MARCELO ROLDAO MATOS · 0 citations
#large language models Open access Aug 2026

Large Language Models in the Acute Stroke Pathway: A Scoping Review of Applications, Evidence Maturity, and Implementation Readiness

Background. Large language models (LLMs) have been rapidly adopted in medicine since late 2022, yet their role in the time-critical acute stroke pathway—from symptom recognition and prehospital triage to emergency diagnosis, imaging-related text tasks, reperfusion decision support, and acute-phase documentation and communication—has not been systematically mapped. Existing reviews cover the whole stroke-care continuum or mix LLMs with traditional NLP, leaving the acute phase under-characterized. Objective. To map the applications, evidence maturity, and implementation readiness of LLMs across the acute stroke pathway. Methods. This scoping review follows the PRISMA-ScR guideline. We search PubMed/MEDLINE, Europe PMC (including preprints), and Google Scholar for studies published from November 2022 onward. Eligible studies center on LLMs/generative AI applied to any stage of the acute stroke pathway. Two reviewers independently screen records and chart data using a piloted form. Evidence is synthesized along two dimensions: five pathway stages (prehospital recognition/dispatch; emergency triage and differential diagnosis; imaging-related text tasks; reperfusion decision support; acute documentation and communication) and three evidence-maturity tiers (simulation/benchmark; retrospective real-world data; prospective deployment). Implementation barriers (hallucination, bias, privacy, regulation, liability, integration, cost) are thematically summarized. Registration note. This review is registered on OSF; the full protocol is available in the attached files.

Xianmu Luo, Hongsong Li, Xiaoli Liao · 0 citations
#large language models Dataset Open access Aug 2026

Reproduction Package - Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection

Measurement harness and results for the paper: Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection Patrick Deininger and Wolfgang Slany. Submitted to MDPI Computers.

Patrick Deininger, Wolfgang Slany · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.