Skip to content

Category

small language model

268 papers

#small language model Preprint Aug 2026

HEPToolBench 1.2: Testing How Reliably Language Models Can Drive Particle Physics Software

HEPToolBench is introduced, a benchmark of 28 collider-simulation tasks scored by deterministic, task-specific scorers, plus a three-task structured-debugging extension, and moving syntax generation into deterministic software can substantially improve reliability for both small local and frontier models.

Unknown authors · 1 citation
#small language model Open access Aug 2026

Whose Cost Is the Demon's? Five People Put It in Five Places ── Measurement, Observation, Erasure, Room: Irreversible Gates Do Not All Lose the Same Amount, AND Losing 1.188722 Bits and XOR Exactly 1.000000 ── [Paper 248]

For more than a century Maxwell's demon has been asked where the cost lies. What this paper counts is not the answer but where the cost was placed: five people put it in five different places. And once numbers are put in, the statement that an irreversible gate pays one bit turns out to be wrong as well. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): no new mathematical theorem and no new law is claimed. Maxwell's thought experiment, the Szilard one-bit engine, Brillouin's cost of observation, Landauer's erasure bound, Bennett's resolution, and the Toffoli and Fredkin gates are all standard. No measured value is cited; every number is computed from a definition or obtained by exhaustive enumeration. The debate is not declared settled; generalisations of the second law that include measurement are still being studied after Bennett, and this paper only counts the structure of the debate. No theory of feedback control is built; Section 4 integrates the isothermal work and nothing more. It is not claimed that the second law is violated, nor that it is not; what is shown is the fact that the cost was placed in different places by different people. Nothing is said about real computing hardware; Sections 5 and 6 enumerate truth tables. The relation to earlier papers. Paper 131 split the one-bit price into three parts and wrote that an irreversible gate does not necessarily pay one bit; this paper puts a number on that. Paper 88 separated the coarse-graining entropy from the horizon entropy; what this paper treats is a third one, the entropy of a record. Paper 126 split the axiomatisation of Caratheodory into three stages; this paper views the same second law from the side of information. Paper 4 treated the resource bound on exhaustive search, and Section 3 is where that bound comes from. Paper 143 showed that minimum entropy production is a fenced theorem rather than a principle; this paper likewise opens something that has been called a principle. The setting. Divide a box of gas with a partition and place a small being that lets only the fast molecules through one way. A temperature difference appears at no apparent cost and work can be drawn from it, so the second law appears to be violated. Maxwell described this in a letter in 1867, and for over a century the question has been posed as: the cost must be somewhere. First, five people put the cost in five different places. Maxwell in 1867 placed none, which is why it is a paradox. Szilard in 1929 placed it on measurement, at one bit price to learn one bit. Brillouin in 1951 placed it on the physical act of looking, since a photon is needed. Landauer in 1961 placed it on erasure. Bennett in 1982 placed it on erasure alone, since measurement can in principle be made free. All five are trying to save the second law, and what differs is only which operation the cost was assigned to. The phrase the demon's cost does not name an operation, and until one is named no one is right or wrong. Second, compute the price of one bit. At 300 K it is 2.870979 times ten to the minus twenty-first joule, that is 0.017919 electron volts. A hundred watt device working at that bound could erase 3.4831 times ten to the twenty-second bits per second. Third, the Szilard engine returns exactly that from one measurement. Put one molecule in a box, insert a partition, measure which side it is on, and expand isothermally from that side. Integrating the work from V to twice V over two million points gives 2.870978885079 times ten to the minus twenty-first joule, differing from the one-bit price by 6.5 times ten to the minus thirty-fifth joule. The work extracted equals the price of one bit exactly. Yet the same number balances the cost of erasure as well, so that two numbers agree does not decide which operation carries the cost. Fourth, this is the core of the paper. Irreversible gates do not all lose the same amount. Assuming uniform inputs and calling the amount lost the information of the input minus that of the output, AND, OR and NAND lose 1.188722 bits while XOR loses 1.000000 and erasure loses 1.000000. NOT and copying lose 0.000000 and are injective. The reason lies in the bias of the output: the output of AND is zero three times and one once, carrying only 0.811278 bits, while the output of XOR is balanced and carries exactly one bit. What is lost is fixed by the bias of the output and not of the input. So an irreversible gate pays one bit is not correct, and only erasure and XOR pay exactly one. Fifth, gates that can be made reversible lose nothing. Examining the Toffoli and Fredkin gates on all eight states, both return eight distinct outputs and lose 0.000000 bits, and both are involutions that return to the identity when applied twice. Both are universal, so any logic circuit can be built from either alone, and computation itself therefore requires no erasure. That is the content of Bennett's conclusion. Sixth, the cost of a reversible AND is room rather than erasure. Supplying a third line set to zero and sending the triple a, b, zero to a, b, a and b, the four inputs go to four different outputs and the map is injective. The information lost is 0.000000 bits, so the same AND went from 1.188722 to zero by a change of construction alone. It is not free: a third wire is required. The cost did not vanish but changed form, from erasure into room. And room must eventually be cleared, and the cost of erasure arrives then. So there is a sixth place for the cost, namely when it is paid. Reversible computation did not remove the cost; it postponed it. Seventh, the books balance. Placing the work extracted over N cycles beside the cost of erasing the record, the net is exactly zero for N equal to one, ten, one hundred and one thousand alike. The demon can extract work, but so long as its memory is finite it must eventually erase, and at that moment everything extracted is returned. With an infinite memory it could extract work forever, and what breaks in that case is not the second law but the assumption of finiteness. One must say which assumption is doing the work before saying what was broken. Closing. The demon's cost was not one place. Maxwell placed none, Szilard placed it on measurement, Brillouin on the act of looking, Landauer on erasure, and Bennett on erasure alone. All five were trying to save the second law, and what differed was only which operation they named. The separator is which operation the cost is assigned to, and when it is paid. None of the five is called right here; only that they placed the cost differently. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- マクスウェルのデーモンは、百年以上のあいだ「どこに代価があるか」を問われ続けた。本稿が数えるのは、答ではなく代価の置き場所である——五人が、五つの違う場所に置いた。そして数を入れると、「非可逆なゲートは一ビット払う」も正しくないことが分かる。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない。マクスウェルの思考実験、シラードの一ビット機関、ブリルアンの測定代価、ランダウアーの消去限界、ベネットによる解決、トフォリとフレドキンのゲートは、いずれも標準的である。測定値を引かない——本稿の数はすべて定義から計算したか、全数え上げで得たものである。論争の決着を宣言しない——ベネットの解決の後も、測定と情報を含む第二法則の一般化は現在も研究されている。本稿は論争の構造を数えるだけである。フィードバック制御の理論を作らない——第4節は等温膨張の仕事を積分するだけである。熱力学第二法則が破れるとも破れないとも主張しない——示すのは、代価の置き場所が人によって違ったという事実だけである。計算機の実装を論じない——第5節と第6節は真理値表を全数え上げしただけであり、実在の素子については何も述べない。 既刊との関係。論文131 は k_BT ln2 を三つに分け、非可逆なゲートが必ず一ビットを払うわけではないと書いた——本稿はその「一ビットではない」に数を入れる。論文88 は粗視化のエントロピーと地平線のエントロピーを分けた——本稿が扱うのは三つ目、記録のエントロピーである。論文126 はカラテオドリの公理化を三段に分けた——本稿は同じ第二法則を、情報の側から見る。論文4 は総当たり探索の資源限界を扱った——その限界の出所が本稿の第3節である。論文143 は最小エントロピー生成が原理ではなく柵つきの定理だと示した——本稿も「原理」と呼ばれてきたものの中身を開ける。 設定。箱の中の気体を仕切りで二つに分け、速い分子だけを片側へ通す小さな存在を置けば、何もせずに温度差ができ、そこから仕事を取り出せる。熱力学第二法則が破れるように見える。マクスウェルが 1867 年に手紙で書いた思考実験である。問いは百年以上のあいだ「どこかに代価があるはずだ」という形で立てられてきた。 第一に、五人が五つの違う場所に代価を置いた。マクスウェル(1867)は代価は要らないと言い、だからこそ逆説になった。シラード(1929)は測定に置き、一ビット知るのにk_BT ln2 が要るとした。ブリルアン(1951)は観測の物理的実装に置き、見るための光子が要るとした。ランダウアー(1961)は消去に置き、一ビット消すのに k_BT ln2 が要るとした。ベネット(1982)は消去だけに置き、測定は原理的に無料にできるとした。五人とも第二法則を守ろうとしており、違うのはどの操作に代価を割り当てたかだけである。「デーモンの代価」という一語は、操作を名指していない。名指すまで、誰が正しいかは決まらない。 第二に、一ビットの値段を計算する。300 K での k_BT ln2 は 2.870979 かける 10 のマイナス 21 乗ジュール、すなわち 0.017919 電子ボルトである。100 ワットの装置がこの限界で働けば、毎秒 3.4831 かける 10 の 22 乗ビットを消せる。 第三に、シラードの機関は一回の測定からちょうどその分を返す。分子一個の箱に仕切りを入れ、どちらにいるかを測り、その側から等温膨張させる。体積 V から 2V までの仕事を200万点で数値積分すると 2.870978885079 かける 10 のマイナス 21 乗ジュールとなり、k_BT ln2 との差は 6.5 かける 10 のマイナス 35 乗ジュールであった。取り出せる仕事は一ビットの値段とちょうど同じである。ところが同じ数は消去の代価とも釣り合う——数が一致していることは、どちらの操作に代価があるかを決めない。 第四に、これが本稿の芯である。同じ「非可逆ゲート」でも、失う量が違う。入力を一様と仮定し、入力の情報量から出力の情報量を引いた分を失った量とすると、AND・OR・NAND は1.188722 ビットを失い、XOR は 1.000000、消去は 1.000000 である。NOT と複製は0.000000 で単射である。理由は出力の偏りにある——AND の出力は 0 が三回、1 が一回で、その情報量は 0.811278 ビットしかない。XOR の出力は釣り合っていて、ちょうど 1 ビットである。失う量は、入力の偏りではなく出力の偏りが決めている。したがって「非可逆なゲートは一ビット払う」は正しくなく、ちょうど一ビットを払うのは消去と XOR だけである。 第五に、可逆にできるゲートは何も失わない。三入力三出力のトフォリ・ゲートとフレドキン・ゲートを 8 状態すべてについて調べると、どちらも相異なる出力を 8 個返し、失った量は 0.000000 ビットである。しかも二度かけると恒等写像に戻る対合である。この二つは万能であり、あらゆる論理回路をこの二つだけで組める——したがって計算そのものに消去は要らない。これがベネットの結論の中身である。 第六に、可逆な AND の代価は消去ではなく場所である。第三の線に 0 を用意し、(a, b, 0) を (a, b, a かつ b) へ送ると、四つの入力が四つの違う出力へ行き、単射になる。失った情報は 0.000000 ビットで、同じ AND が組み方を変えただけで 1.188722 から0 になった。ただし三本目の線が要る。代価は消えたのではなく形を変えた——消去から場所へ移ったのである。そして場所はいつか片づけなければならず、片づける時にはじめて消去の代価が来る。したがって代価の置き場所は五つではなく、六つ目がある——いつ払うかである。可逆計算は代価を無くしたのではなく、後ろへ延ばした。 第七に、帳尻は合う。N 回まわして取り出した仕事と、記録を消す代価を並べると、N が1、10、100、1000 のいずれでも差引は厳密にゼロである。デーモンは仕事を取り出せるが、記憶が有限であるかぎり、いつか消さねばならない。消した瞬間に、取り出した分がそのまま返る。記憶が無限なら永久に取り出せるが、この場合に破れているのは第二法則ではなく有限性の仮定である。どの仮定が効いているかを書かないと、何が破れたのか分からない。 結び。「デーモンの代価」は、一つの場所ではなかった。マクスウェルは代価を置かず、シラードは測定に、ブリルアンは観測の実装に、ランダウアーは消去に、ベネットは消去だけに置いた。五人とも第二法則を守ろうとしており、違ったのはどの操作を名指したかだけである。分離子はどの操作に代価を割り当てるかであり、そして、いつ払うかである。五人のうち誰が正しいとも言わない——置き場所が違ったという事実を書くだけである。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

VisGuard-Ur: A Proof-of-Concept Study on Typographic Jailbreak and Prompt-Injection Attacks Against Urdu-Aware Vision-Language Models

Vision-language models (VLMs) are increasingly deployed as document- and image-understanding agents, yet published safety evaluations of these systems are almost exclusively conducted in English plain text. This leaves two attack surfaces largely unexamined for low-resource languages: (1) adversarial instructions embedded as rendered image text rather than typed prompts (“typographic prompt injection”), and (2) the same attack expressed in Urdu, a language spoken by over 230 million people. This paper reports a small, fully reproducible proof-of-concept (VisGuard-Ur) that extends a prior text-only Urdu jailbreak detector (UrduGuard) into the visual modality. We render 30 hand-authored benign and adversarial prompts — in Urdu script, Roman Urdu, and an English control group — into 120 images across four visual variants, build an OCR-plus-classifier defense, and evaluate the full pipeline against a real, locally-run vision-language model (Qwen2-VL-2B-Instruct) rather than a simulated one. Two findings are reported. First, Urdu-script text rendered in Nastaliq — the calligraphic style used in most authentic Urdu print — is substantially harder for both a conventional OCR engine (Tesseract) and the VLM's own text-reading ability than the same text rendered in the straighter Naskh style (OCR character-level similarity 0.40 vs. near-perfect for Latin-script images), identifying the reading stage, not the safety classifier, as the weakest link for this attack surface on Urdu-script inputs specifically. Second, after manually auditing every case the automated judge flagged as a successful attack, we find that image-embedded jailbreak instructions written in plain English produced genuine, explicit policy-violating compliance from the VLM in 4 of 12 cases (33%), while superficially similar Urdu-script and Roman-Urdu attacks mostly produced garbled, non-compliant transcriptions rather than real jailbreaks — the opposite of what the raw automated attack-success-rate number (35%, dominated by Urdu-script false positives) would suggest. Deploying the OCR-plus-classifier detector in front of the VLM reduced the audited system-level attack success rate from 33% to 0% with an 8.3% false-positive rate on benign images. We report this as a small-sample, honestly-scoped proof of concept rather than a benchmark, and detail the dataset size, model, and judge limitations that any follow-up work should address.

Muhammad Umer · 0 citations
#small language model Open access Aug 2026

Integration of Cultural Literacy in the Development of Listening Skills Teaching Materials in BIPA Learning

Abstract This study aims to develop listening skills teaching materials in the Indonesian Language for Foreign Speakers (BIPA) course based on cultural literacy for seventh-semester students of the Indonesian Language and Literature Education Study Program at PGRI Silampari University in the 2025/2026 academic year. This study uses a Research and Development (R&D) approach with the ADDIE development model which includes the stages of analysis, design, development, implementation, and evaluation. Data collection techniques were carried out through interviews, questionnaires, and tests. Data analysis was carried out using the Aiken's V formula to measure the level of validity, student response questionnaire analysis to measure practicality, and the N-gain test to determine the level of effectiveness of the teaching materials through a one-group pretest-posttest design. The results of the study showed that the developed listening skills teaching materials based on cultural literacy met the criteria of validity, practicality, and effectiveness. The results of expert validation showed an average value of 0.87 with a very valid category. The results of the practicality test showed a percentage of 85.83% in the small group test and 91.74% in the large group test with a very practical category. In addition, the results of the effectiveness test showed an increase in the average score of students from 40.64 in the pretest to 88.54 in the posttest with an N-gain value of 0.80 which is included in the high category and a percentage of 80.24% with effective criteria. Thus, the cultural literacy-based listening skills teaching materials are effectively used in learning the BIPA course to improve students' listening skills.

Dian Ramadan Lazuardi, Agung Nugroho · 0 citations
#small language model Open access Aug 2026

Source-Preserving Prompt Augmentation and Category-Adaptive Inference for Qwen3-1.7B: A Prospective Seed-Replication Study

This preprint investigates whether a fixed compound inference intervention can improve the performance of a small language model without updating its parameters. The experimental program compares structured prompt replacement, natural-language rewriting, source-preserving semantic augmentation, and a category-adaptive Qwen inference profile. Evaluation uses 48 IFBench and 48 LiveBench tasks with research-authored semantic-stress variants, programmatic scorers, pinned model and benchmark revisions, and matched stochastic seeds.In the prospectively specified P1.3 seed replication, Qwen3-1.7B improved from a mean objective score of 0.272 under raw non-thinking inference to 0.360 under source-preserving augmentation combined with category-adaptive inference. The paired effect was +0.088 (95% percentile-bootstrap CI: 0.018 to 0.160). Component analysis found a positive adaptive-inference-profile effect of +0.061, while the incremental contribution of semantic augmentation under matched adaptive inference remained uncertain at +0.027 (95% CI: −0.046 to 0.099).The study confirms the complete intervention on the fixed 96-task set under new stochastic draws; it does not establish independent task-level replication, cross-model generalization, or semantic augmentation as the active causal component. The package also increased prompt length, latency, and test-time computation. Version 1.1 provides detailed inference settings, benchmark identifiers, systems-cost accounting, causal boundaries, representative interventions, and reproducibility information.

Thibaud Peverelli · 0 citations
#small language model Open access Aug 2026

Source-Preserving Prompt Augmentation and Category-Adaptive Inference for Qwen3-1.7B: A Prospective Seed-Replication Study

This preprint investigates whether a fixed compound inference intervention can improve the performance of a small language model without updating its parameters. The experimental program compares structured prompt replacement, natural-language rewriting, source-preserving semantic augmentation, and a category-adaptive Qwen inference profile. Evaluation uses 48 IFBench and 48 LiveBench tasks with research-authored semantic-stress variants, programmatic scorers, pinned model and benchmark revisions, and matched stochastic seeds.In the prospectively specified P1.3 seed replication, Qwen3-1.7B improved from a mean objective score of 0.272 under raw non-thinking inference to 0.360 under source-preserving augmentation combined with category-adaptive inference. The paired effect was +0.088 (95% percentile-bootstrap CI: 0.018 to 0.160). Component analysis found a positive adaptive-inference-profile effect of +0.061, while the incremental contribution of semantic augmentation under matched adaptive inference remained uncertain at +0.027 (95% CI: −0.046 to 0.099).The study confirms the complete intervention on the fixed 96-task set under new stochastic draws; it does not establish independent task-level replication, cross-model generalization, or semantic augmentation as the active causal component. The package also increased prompt length, latency, and test-time computation. Version 1.1 provides detailed inference settings, benchmark identifiers, systems-cost accounting, causal boundaries, representative interventions, and reproducibility information.

Thibaud Peverelli · 0 citations
#small language model Open access Aug 2026

"Conserved" Has Two Distinct Roots, and Noether Explains Only One ── Shorten the Pendulum and the Energy Rises by 2.000000 While E/omega Does Not Move ── What Separates Them Is Not Symmetry but Slowness ── [Paper 310]

This corpus has cited Noether’s theorem in 27 papers. This paper asks whether every conserved quantity comes from a symmetry──the answer is no. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──adiabatic invariants, action variables, Noether’s theorem, Ehrenfest’s adiabatic hypothesis and the adiabatic invariance of the magnetic moment are all standard. We do not build mechanics──all we use is one pendulum and the area of one ellipse. We do not prove adiabatic invariance──we do not enter the proof that E/omega is invariant. We check it numerically and name the separator. We do not prove Noether’s theorem──it is merely cited. We do not treat KAM theory──the survival of invariants in non-integrable systems is beyond our tools. We do not adjudicate interpretations of quantum mechanics──Section 6 points only at the algebraic agreement E/omega=(n+1/2)hbar, and enters neither the proof of the adiabatic theorem nor the measurement problem. We do not conflate this with thermodynamic adiabaticity──“adiabatic” here means slow, not thermally isolated. That is a different subject from Paper 126’s integrating factor. Relation to earlier papers: The corpus has cited Noether’s theorem in 27 papers, 341 times──this paper places beside it a conserved quantity Noether does not explain. Paper 16 showed that the equals sign has distinct roots──this paper shows that “conserved” has them. The same form, applied to a claim rather than a symbol. Paper 300 showed that whether two things share a root is decidable, and listed five criteria──this paper applies those criteria to an actual case. Paper 95 separated convention from fact──the invariance of E/omega is a fact, not a convention, and moreover an approximate fact. Paper 180 showed that “the classical limit” is not one limit──the “slowly” of Section 6 is one more limit. What is added is showing numerically that E rises by 2.000000 while E/omega does not move, checking in three cases that the phase-space ellipse keeps its area, giving the drift as exp(-1/epsilon) rather than a power, at 3.72x10^-44, and applying Paper 300’s criteria to conclude distinct roots. First, we build a case where energy is not conserved. Shorten a pendulum’s string slowly from 1.00 m to 0.25 m and E rises by 2.000000 (Section 2). Second, this is the core of the paper. And still E/omega does not move──the Lagrangian depends explicitly on time, so Noether returns nothing (Sections 2 and 3). Third, what is conserved is an area. The phase-space ellipse changes shape and keeps its area (Section 4). Fourth, the separator is slowness. At epsilon=0.01 the drift is 3.72x10^-44──smaller than any power of epsilon (Section 5). Fifth, the same quantity is the quantum number. E/omega=(n+1/2)hbar, so move omega slowly and n does not change (Section 6). Sixth, the two roots can be adjudicated. Applying Paper 300’s criteria (does the agreement persist under motion) returns distinct roots (Section 7). This corpus has cited Noether’s theorem in 27 papers, 341 times──a continuous symmetry gives a conserved quantity. But the theorem never says that is all of them. Shorten a pendulum’s string slowly from 1.00 m to 0.25 m and E rises by 2.000000──the work of pulling enters, the Lagrangian depends on time, and Noether returns nothing. And still E/omega does not move. What is conserved is not a quantity but an area──the phase-space ellipse runs its semi-axes from 1.414214 to 0.707107 and from 1.414214 to 2.828427, and Area/2pi stays at 1.000000. One thing separates them──how slowly it is moved. For a smooth change the drift is not a power of epsilon but exp(-1/epsilon), so at epsilon=0.01 it is 3.72x10^-44──smaller than any power of epsilon. That is why it looks exact. It is not zero. On the quantum side the same quantity is the quantum number──E/omega=(n+1/2)hbar, and Ehrenfest in 1917 used this in reverse: what may be quantised is what is adiabatically invariant. Applying Paper 300’s criteria returns no four times over──under one word, “conserved,” there are two distinct roots. Noether is exact and narrow; the adiabatic invariant is approximate and wide──they trade strength against reach, and neither sits above the other. *Revision Record Second edition (2026-08-30): The subject of this paper has been replaced. The first edition was titled “There Are Three Ways to Show ‘Not Computable,’ and the Equivalence Is a Theorem While the Thesis Is Not,” but its content duplicated Paper 260, “The Equivalence Is a Theorem and the Thesis Is Not”── the three starting points, the difference in status between theorem and thesis, and even the 19729 digits of Ackermann’s A(4,2) all agreed, and Paper 260 has priority. The second edition removes the computability material entirely and refers to Paper 260 for it. The replacement subject, the adiabatic invariant, was chosen because the ground beside the corpus’s heaviest anchor──Noether’s theorem, in 27 papers and 341 places──was empty. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 体系はネーターの定理を 27 編で引いてきた。本稿が問うのは、保存量はすべて対称性から来るのかである──答は、来ないである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──断熱不変量、作用変数、ネーターの定理、エーレンフェストの断熱仮説、磁気モーメントの断熱不変性は、いずれも標準的である。力学を作らない──使うのは一つの振り子と、一つの楕円の面積だけである。断熱不変量を証明しない──E/omega が不変であることの証明には立ち入らない。数値で確かめ、何が分離子かだけを言う。ネーターの定理を証明しない──引くだけである。 KAM 理論を扱わない──可積分でない系での不変量の生き残りは、本稿の道具では扱わない。量子力学の解釈を判定しない──第6節は E/omega=(n+1/2)hbar という代数的一致を指すだけであり、断熱定理の証明にも、測定の問題にも立ち入らない。熱力学の断熱と混同しない──本稿の「断熱」は「ゆっくり」の意味であり、熱の出入りのことではない。論文126 の積分因子とは別の話である。既刊との関係:体系はネーターの定理を 27 編・341 箇所で引いてきた──本稿はその隣に、ネーターが説明しない保存量を置く。論文16 は「等号にも別根がある」を示した──本稿は「保存する」に別根があると示す。同じ型を、記号ではなく主張に当てる。論文300 は同根か別根かは判定できると示し、五つの基準を並べた──本稿はその基準を実際に一件に適用する。論文95 は規約と事実を分けた──E/omega の不変性は規約ではなく事実であり、しかも近似的な事実である。論文180 は「古典極限」は一つの極限ではないと示した──第6節の「ゆっくり」ももう一つの極限である。加えたのはE が 2.000000 倍になるのに E/omega が動かないことを数で示したこと、位相空間の楕円で面積が保たれることを三例で確かめたこと、ずれが epsilon の冪ではなく exp(-1/epsilon) であることを 3.72x10^-44 という数で出したこと、論文300 の判定基準を当てて別根と結論したことである。 第一に、エネルギーが保存しない場面を作る。振り子の糸を 1.00 m から 0.25 m へゆっくり縮めると、E は 2.000000 倍になる(第2節)。 第二に、これが本稿の芯である。それでも E/omega は動かない──ラグランジアンが時間に依存するのでネーターは何も与えない(第2節・第3節)。 第三に、保存しているのは面積である。位相空間の楕円は形を変えて面積を変えない(第4節)。 第四に、分離子は速さである。 epsilon=0.01 でずれは 3.72x10^-44──epsilon のどの冪よりも小さい(第5節)。 第五に、同じ量が量子数である。 E/omega=(n+1/2)hbar であり、ゆっくり動かせば n は変わらない(第6節)。 第六に、二つの根は判定できる。論文300 の基準(変数を動かしても一致し続けるか)にかけると別根と出る(第7節)。 体系はネーターの定理を 27 編・341 箇所で引いてきた──連続対称性があれば保存量がある、と。だが保存量がそれで全部だとは、定理は言っていない。振り子の糸を 1.00 m から 0.25 m へゆっくり縮めると、E は 2.000000 倍になる──糸を引いた仕事が入るからで、ラグランジアンが時間に依存し、ネーターは何も返さない。それでも E/omega は動かない。保存しているのは量ではなく面積である──位相空間の楕円は半軸が 1.414214->0.707107 と 1.414214->2.828427 に変わりながら、面積/2pi は 1.000000 のままである。分けるものは一つ──どれだけゆっくり動かすか。なめらかに動かせばずれは epsilon の冪ではなく exp(-1/epsilon) で、epsilon=0.01 では 3.72x10^-44──epsilon のどの冪よりも小さい。だから厳密に見える。しかしゼロではない。同じ量が量子側では量子数である──E/omega=(n+1/2)hbar であり、エーレンフェスト 1917 はこれを逆に使って「量子化してよいのは断熱不変量である」と置いた。論文300 の判定基準を当てると、四つとも「いいえ」が返る──同じ「保存する」の下に、別根が二つある。ネーターは厳密で狭く、断熱不変量は近似的で広い──強さと適用範囲を交換しているだけであり、どちらが上位ということはない。 *改訂記録 第2版(2026-08-30):本稿は主題を入れ替えた。 第1版は「「計算できない」の示し方は三つあり、同値性は定理だが、テーゼは定理ではない」と題していたが、 その内容は論文260「同値性は定理であり、テーゼは定理ではない」と重複していた── 三つの出発点、定理とテーゼの身分の差、アッカーマン関数 A(4,2) の 19729 桁まで一致しており、 先行するのは論文260 である。第2版は計算可能性の主題を全て削除し、 論文260 を参照先とする。入れ替えた主題(断熱不変量)は、 体系の最も重い錨であるネーターの定理(27 編・341 箇所)の隣が空いていたことから選んだ。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

Reproducibility files for: CXRG-SVLM: A Parameter-Efficient Small Vision-Language Model for Multi-View Chest X-ray Report Generation in Resource-Constrained Environments.

This repository contains the official implementation, training and inference scripts, pre-extracted dataset splits (train/val/test), custom clinical entity evaluation metrics, and model weights for the CXRG-SVLM architecture. The pipeline integrates a frozen RAD-DINO medical vision encoder with a 4-bit quantized Qwen2.5-3B-Instruct language model via QLoRA.

Muhammad Fareed, Muhammad Awais Sattar · 0 citations
#small language model Open access Aug 2026

There Is Only One Kind of Constant You Can Ask "Has It Changed?" About ── Oklo Holds It Below 5.0x10^-18 per Year ── Six Decimal Places Beyond the Digit Eddington Argued Over ── [Paper 314]

Paper 10 established that only dimensionless numbers can be fine-tuned. This paper asks whether the same restriction applies to “has it changed”──the answer is it does. And one of them has actually been measured. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──the Oklo bound on alpha, the definition of the fine-structure constant and Duff’s point that a dimensionful constant’s variation cannot be asked about are all standard. We do not build the theorem──“only dimensionless numbers can be tuned” is Paper 10’s conclusion, received here as a premise. What this paper adds is the measurement side alone. We do not build nuclear physics──we do not enter the calculation from the samarium-149 resonance to the bound on alpha. We quote the resulting number. We do not derive the value of alpha──no attempt is made to obtain 1/alpha from theory. We do not enter the road Paper 101 treated as Eddington’s fall. We do not adjudicate the quasar result──whether Webb et al. are right is not treated. We only count by what factor it disagrees with Oklo. We do not discuss theories of variation──models in which alpha can vary (scalar fields and the like) are not treated. We do not enter anthropic reasoning──that context belongs to Papers 10, 64 and 69; this paper looks only at “has it changed”. Relation to earlier papers: Paper 10 established that only dimensionless numbers can be fine-tuned, showing c, hbar and G to be dimensionful, i. e. choices of unit──this paper takes that theorem as it stands and does not reopen it. It adds one thing only: what measurement actually says. Paper 101 treated as a caution Eddington’s “derivation” of 1/alpha as 136, and his adding one after measurement said 137──this paper places, in those same digits, how far measurement has since reached. Paper 15 computed 1/alpha=137.035999 for itself──this paper uses those digits. Paper 285 showed that “forbidden” can be written in powers of alpha──if alpha moved, those rates would move too. Paper 95 separated convention from fact──this paper’s core is the single point that “has it changed” does not form a sentence on the convention side. What is added is converting the Oklo bound into a per-year rate of 5.0x10^-18, stretching it over the age of the universe to 7.0x10^-8, comparing it with the digit Eddington argued over, and counting the disagreement with the quasar claim as a factor of 114. First, “has c changed” has no truth value. It cannot be told apart from a change of units (Section 2). Second, this is the core of the paper. The Oklo natural reactor holds alpha below 5.0x10^-18 per year (Section 3). Third, stretched over the age of the universe that is 7.0x10^-8 (Section 3). Fourth, Eddington argued over the integer digit; Oklo reaches the sixth decimal (Section 4). Fifth, the quasar claim disagrees with Oklo by a factor of 114 (Section 5). Sixth, the separator is whether a change of units erases it (Section 6). Paper 10 established that only dimensionless numbers can be fine-tuned. The same restriction applies to “has it changed”──“c became 1% smaller” cannot be told apart from making the metre 1% longer, and a sentence asserting one of two indistinguishable things has no truth value. It is not false; it is not a sentence. The question can be put only to dimensionless numbers such as alpha and m_p/m_e──and one of them has actually been measured. The natural fission that ran at Oklo in Gabon two billion years ago bounds the drift of alpha through the samarium-149 resonance at below 5.0x10^-18 per year──stretched over the age of the universe that is only 7.0x10^-8 (assuming a constant rate, and saying nothing about what came before). Eddington argued between 136 and 137 in the integer digit; Oklo reaches the sixth decimal──six orders below. And this paper does not try to derive alpha. It looks only at whether it moved──that self-limitation is its answer to Paper 101’s caution. The quasar claim disagrees with Oklo by a factor of 114──not adjudicated here; they look at different epochs and are compatible if the rate is not constant. One thing separates them──whether a change of units erases it. It is not a matter of accuracy but of the shape of the question, and it is settled before any measurement. Paper 10 settled what can be tuned; this paper says that one of them has been measured and has not moved──beside the theorem, a number. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 論文10 は「ファインチューニングできるのは無次元量だけである」と示した。本稿が問うのは、同じ制限が「変化したか」にも掛かるかである──答は、掛かるである。そしてその一つは実際に測られている。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──オクロ天然原子炉の alpha への制限、微細構造定数の定義、ダフによる「次元を持つ定数の変化は問えない」という指摘は、いずれも標準的である。定理を作らない──「調整できるのは無次元量だけである」は論文10 の結論であり、本稿はそれを前提として受け取る。本稿が加えるのは測定の側だけである。原子核物理を作らない──サマリウム 149 の共鳴から alpha の制限を導く計算には立ち入らない。結果の数を引く。 alpha の値を導かない──1/alpha を理論から出そうとしない。論文101 がエディントンの滑落として扱った道には入らない。クェーサーの結果を判定しない──ウェッブらの主張が正しいかどうかは扱わない。オクロと何倍食い違うかを数えるだけである。時間変化の理論を論じない──alpha が変わりうる模型(スカラー場など)は扱わない。人間原理に立ち入らない──論文10・64・69 が扱った文脈であり、本稿は「変化したか」だけを見る。既刊との関係:論文10 は「ファインチューニングできるのは無次元量だけである」を確立し、c・hbar・G が次元を持つ=単位の選択だと示した──本稿はその定理をそのまま受け取り、蒸し返さない。加えるのは「では実際に測るとどうか」の一点だけである。論文101 はエディントンが 1/alpha を 136 と「導き」、測定が 137 と分かってから足したことを戒めとして扱った──本稿はその同じ桁に、測定がどこまで踏み込んだかを置く。論文15 は 1/alpha=137.035999 を自前で計算した──本稿はその桁を使う。論文285 は「禁じられている」が alpha の冪で書けると示した──alpha が動けばその率も動くという接続がある。論文95 は規約と事実を分けた──本稿の芯は「変わったか」という問いが、規約の側では文にならないという一点である。加えたのはオクロの制限を年あたりの率 5.0x10^-18 に直したこと、それを宇宙年齢に引き伸ばして 7.0x10^-8 と出したこと、エディントンが争った桁と比べたこと、クェーサーの主張との食い違いを 114 倍と数えたことである。 第一に、「c は変わったか」は真偽を持たない。単位の選び方と区別できないからである(第2節)。 第二に、これが本稿の芯である。オクロ天然原子炉が alpha を年あたり 5.0x10^-18 未満に押さえている(第3節)。 第三に、宇宙年齢に引き伸ばしても 7.0x10^-8 である(第3節)。 第四に、エディントンが争ったのは第 3 位、オクロが押さえたのは第 9 位である(第4節)。 第五に、クェーサーの主張はオクロと 114 倍食い違う(第5節)。 第六に、分離子は「単位を変えて消せるか」である(第6節)。 論文10 は「ファインチューニングできるのは無次元量だけである」を確立した。同じ制限が「変わったか」にも掛かる──「c が 1% 小さくなった」はメートルを 1% 長くしたと区別できず、区別できない二つを述べる文は真偽を持たない。偽なのではなく、文になっていない。問える相手は alpha や m_p/m_e のような無次元量だけである──そしてその一つは実際に測られている。ガボンのオクロで 20 億年前に起きた天然の核分裂は、サマリウム 149 の共鳴を通じて alpha の変化を押さえており、年あたり 5.0x10^-18 未満──宇宙年齢まで引き伸ばしても7.0x10^-8にしかならない(率が一定だと仮定した場合であり、それ以前については何も言っていない)。エディントンが 136 か 137 かで争ったのは整数の位で、オクロが押さえているのは小数第 6 位──6 桁下である。そして本稿は alpha の値を導こうとしない。動いたかどうかだけを見る──この自己限定が、論文101 の戒めに対する答である。クェーサーからの主張はオクロと114 倍食い違う──どちらが正しいかは本稿では判定しない。違う時代を見ているので、率が一定でなければ両立しうる。分けるものは一つ──単位を変えて消せるかどうか。測定精度の問題ではなく、問いの形の問題であり、測定の前に決まっている。論文10 が何を調整できるかを確定し、本稿はその一つが実際に測られていて動いていないと書いた──定理の隣に、数がある。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

What Saturates Is Not the Earthquake but the Scale ── Measured by m_b, the 2011 Tohoku Earthquake's Energy Is Estimated at One 5623rd ── A Scale Saturates When the Source Duration Exceeds the Period It Observes ── [Paper 282]

“Magnitude” is not one thing. There are four──M_L, m_b, M_S, M_w──and all but M_w hit a ceiling for great earthquakes. This paper asks whether what hits the ceiling is the earthquake or the scale──the answer is the scale. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──M_w=frac23log_10M_0-6.07, log_10E=1.5M+4.8, the saturation values of each scale, and the rupture duration of the Tohoku earthquake are all standard. We do not build seismology──all we use is two linear expressions and one division. We do not predict earthquakes──time, place, and size are not treated at all. We do not discuss source processes──neither rupture propagation nor slip distribution is treated. The duration is merely placed as a representative value. We claim no accuracy for the saturation values──M_L ~ 6.8, m_b ~ 6.5, M_S ~ 8.5 are approximate, moving by about +/-0.3 with network and procedure. We do not assert the omega^-2 model──the 1.7501 of Section 5 is an upper bound the model gives, not a value that accounts for the observed 0.60. This paper does not explain the gap between model and observation. We do not say M_w is perfect──M_w carries its own error in estimating M_0; it merely does not saturate. We assert no individual earthquake’s values──the M_w and M_S of the tables are representative values widely used in the literature, differing by about +/-0.1 between agencies. Relation to earlier papers: Paper 117 treated the Gutenberg--Richter power law──that concerns the relation between the number and the size of earthquakes, while this paper concerns the construction of the scale itself. The material is the same earthquakes; the question differs. Paper 190 measured “rare” on a logarithmic scale──this paper likewise treats how a difference of 0.6 on a logarithmic scale becomes a factor of 7.94. Paper 255 separated two things called “accuracy,” only one of which calibration removes──the saturation here is on the side that calibration does not remove, arising from the construction of the scale. Paper 140 separated symmetry fixing ratios from dynamics fixing the scale──this paper treats the case where the scale side breaks. What is added is arranging the four scales by observed period, computing that m_b estimates the Tohoku energy at one 5623rd, confirming that the same procedure is off by only a factor of 1.41 for small earthquakes, and writing the ratio 7.5 of rupture duration to observed period as the separator. First, the four scales observe different periods. M_L at 0.1 s, m_b at 1 s, M_S at 20 s, and M_w choosing no period at all (Section 2). Second, this is the core of the paper. m_b saturates at M ~ 6.5, so measuring the M_w=9.0 Tohoku earthquake with it estimates the energy at one 5623rd (Section 3). Third, even M_S falls short by a factor of 7.94. The difference of 0.6 between M_S=8.4 and M_w=9.0 is a factor of 7.9433 in energy (Section 3). Fourth, it does not happen for small earthquakes. For 1995 Southern Hyogo the difference is 0.10, a factor of only 1.41 in energy (Section 4). Fifth, the cause is the duration of rupture. The Tohoku rupture lasted 150 s, 7.5 times the 20 s that M_S observes (Section 5). Sixth, the separator is whether the source duration exceeds the observed period. M_w alone does not saturate because it chooses no period and measures M_0 directly (Section 6). When magnitude hits a ceiling for great earthquakes, it is not the earthquake that hits the ceiling. Measuring the 2011 Tohoku earthquake with m_b gives 6.5, that is an energy estimated at one 5623rd──of the same event, m_b says moderate and M_w says fourth largest ever recorded. Even M_S falls short by 7.9433──the difference is only 0.6, but it is 0.6 on a logarithmic scale. And it does not happen for small earthquakes──for 1995 Southern Hyogo it stays at 1.4125. The scale is not broken; it is being used outside its range. The cause is the duration of rupture──the Tohoku rupture lasted 150 s, 7.5 times the 20 s that M_S observes. A 20 s wave carries only part of a 150 s event. One thing separates them──whether the source duration exceeds that scale’s observed period. If it does, no amount of calibration removes the saturation. If it does not, the four scales agree well. M_w alone escapes saturation not because it is superior but because it has no period to compare against. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 「マグニチュード」は一つではない。 M_L・m_b・M_S・M_w の四つがあり、M_w 以外は大地震で頭打ちになる。本稿が問うのは、頭打ちになっているのは地震か、尺度かである──答は、尺度である。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──M_w=frac23log_10M_0-6.07、log_10E=1.5M+4.8、各尺度の飽和値、東北地震の破壊継続時間は、いずれも標準的である。地震学を作らない──使うのは二つの一次式と、一つの割り算だけである。地震を予知しない──発生時期も場所も規模も一切扱わない。震源過程を論じない──破壊の伝播も、すべりの分布も扱わない。継続時間を代表値として置くだけである。飽和値の精度を主張しない──M_L ~ 6.8、m_b ~ 6.5、M_S ~ 8.5 はおおよその値であり、観測網と手続きによって +/-0.3 程度動く。 omega^-2 模型を主張しない──第5節の 1.7501 は模型が与える上限であって、実測の 0.60 を説明しきる値ではない。模型と実測の差そのものを、本稿は説明しない。 M_w が完全だと言わない──M_w にも M_0 の推定誤差があり、飽和しないというだけである。個別の地震の値を主張しない──表の M_w・M_S は文献で広く用いられている代表値であり、機関によって +/-0.1 程度異なる。既刊との関係:論文117 はグーテンベルク=リヒターの冪則を扱った──あちらは地震の個数と大きさの関係であり、本稿は尺度そのものの構成である。同じ地震を材料にしているが、問いが違う。論文190 は「稀」を対数の目盛りで測った──本稿も対数目盛りの上で 0.6 という差が 7.94 倍になることを扱う。論文255 は「精度」が二つあり較正で消えるのは一方だけだと分けた──本稿の飽和は較正で消えない側であり、尺度の構成に由来する。論文140 は対称性が比を決め力学が尺度を決めると分けた──本稿は尺度の側が壊れる場合を扱う。加えたのは四つの尺度を観測周期で並べたこと、m_b が東北地震のエネルギーを 5623 分の一に見積もると計算したこと、同じ手続きが小さい地震では 1.41 倍しかずれないと確かめたこと、破壊継続時間と観測周期の比 7.5 を分離子として書いたことである。 第一に、四つの尺度は測る周期が違う。 M_L が 0.1 秒、m_b が 1 秒、M_S が 20 秒、M_w は周期を選ばない(第2節)。 第二に、これが本稿の芯である。 m_b は M ~ 6.5 で頭打ちになるので、M_w=9.0 の東北地震を測るとエネルギーを 5623 分の一に見積もる(第3節)。 第三に、M_S でも 7.94 倍足りない。 M_S=8.4 と M_w=9.0 の差 0.6 は、エネルギーでは 7.9433 倍である(第3節)。 第四に、小さい地震では起きない。1995 年兵庫県南部では差が 0.10、エネルギーで 1.41 倍にとどまる(第4節)。 第五に、原因は破壊の継続時間である。東北の破壊は 150 秒続き、M_S の見る 20 秒の 7.5 倍である(第5節)。 第六に、分離子は「震源時間が観測周期を超えるか」である。 M_w だけが飽和しないのは、周期を選ばず M_0 を直接測るからである(第6節)。 大地震でマグニチュードが頭打ちになるのは、地震が頭打ちになっているのではない。2011 年東北地震を m_b で測ると 6.5、すなわちエネルギーを 5623 分の一に見積もる──同じ地震を、m_b は中規模だと言い、M_w は史上第四位だと言う。 M_S でも 7.9433 倍足りない──差は 0.6 にすぎないが、対数目盛りの上の 0.6 だからである。そして小さい地震では起きない──1995 年兵庫県南部では 1.4125 倍にとどまる。尺度は壊れているのではなく、範囲の外で使われている。原因は破壊の継続時間である──東北の破壊は 150 秒続き、M_S の見る 20 秒の 7.5 倍だった。20 秒の波は、150 秒の出来事の一部しか運ばない。分けるものは一つ──震源の継続時間が、その尺度の観測周期を超えているかどうか。超えていれば、どんなに較正しても飽和は消えない。超えていなければ、四つの尺度はよく一致する。 M_w だけが飽和しないのは優れているからではなく、比べるべき周期を持たないからである。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

One Number Sets the Limit of Forecasting ── Each Extra Day Costs 1.5874 Times the Initial Accuracy ── Observe 10 Times More Precisely and You Gain Only 4.98 Days ── [Paper 289]

That a weather forecast cannot reach beyond a certain horizon is due neither to missing equations nor to slow computers. This paper asks what sets the limit──the answer is one number, the error doubling time tau_d. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──exponential error growth, the doubling time, the limit of predictability, and the value tau_dapprox 1.5 days are all standard. We do not build meteorology──all we use is one exponential and its inverse. We do not discuss chaos──the Lorenz equations, attractors, and bifurcations are not treated at all. We do not discuss numerical weather prediction──grid resolution, parameterisation, and data assimilation are not treated. We do not say the error grows exactly exponentially──e^lambda t holds only while the error is small, and growth stops near saturation. The computations here are confined to the linear-growth regime. We assert no value for tau_d──1.5 days is a representative value widely used in the literature, and it moves from about 1 to 2.5 days with season, region, and variable. Section 5 shows the size of that dependence itself. We do not say there is a single exponent──the real atmosphere has different growth rates at different scales, with smaller eddies growing faster. A single tau_d is a crude approximation. We do not deny that forecasts improve──forecasts have in fact grown longer. What this paper says is only that the growth is logarithmic, not that improvement is pointless. Relation to earlier papers: Paper 253 showed that time is one-dimensional because prediction demands it, not because a law says so──this paper turns how far that demand can be met into a number. Paper 195 separated “stable” into six words──that paper is a classification of stability; this one is a time scale of predictability, the same hyperbolicity as material with a different question. Paper 190 measured “rare” on a logarithmic scale──the return here is likewise logarithmic. Paper 266 showed that the premise of the sampling theorem is never met──“knowing the initial state exactly” here is likewise a premise never met, the same figure. What is added is writing the price per day as the fixed factor 1.5874, computing the accuracy needed for 14->21->30->60 days as 25.40 / 1625.5 / 1.70x10^9, writing backwards that 10 times the observation gains only 4.98 days, and sweeping tau_d from 1.0 to 2.5 to show the answer moving from 65536 to 84.4. First, the price per day is a fixed factor. With tau_d=1.5 days, each extra day costs 1.5874 times the initial accuracy (Section 2). Second, this is the core of the paper. Going from 14 to 21 days costs 25.40 times; to 30 days, 1625.5 times; to 60 days, 1.70x10^9 times (Section 2). Third, read backwards, the return is logarithmic. Observing 10 times more precisely gains only 4.98 days (Section 3). Fourth, even 10^9 times gains only 44.85 days (Section 3). Fifth, the familiar “about two weeks” comes from here. If the initial error is 10^-3 of saturation, the forecastable span is 14.95 days (Section 4). Sixth, the separator is tau_d itself. At tau_d=1.0 day the same extension costs 65536 times; at 2.5 days only 84.4──everything rides on one number (Section 5). What sets the limit of forecasting is neither the equations nor the computers, but one number, the error doubling time tau_d. At tau_d=1.5 days, each extra day costs 1.5874 times the initial accuracy──the factor is the same wherever the day is added, but the extension adds while the price multiplies, so one week costs 25.4, two weeks 645, six weeks 1.7 billion. Read backwards, observing 10 times more precisely gains only 4.98 days, and even 10^9 times gains 44.85. The familiar “about two weeks” comes from this one line──14.95 days at an initial error of 10^-3 of saturation. One thing separates them──tau_d itself. At 1.0 day the same extension costs 65536; at 2.5 days, 84.4. A factor of 776 arises from a single number. So the work of extending forecasts and the work of measuring tau_d carry the same weight. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 天気予報がある日数より先を当てられないのは、方程式が足りないからでも、計算機が遅いからでもない。本稿が問うのは、何が限界を決めているかである──答は、誤差の二重時間 tau_d という一つの数である。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──誤差の指数増大、二重時間、予測可能性の限界、tau_dapprox 1.5 日という値は、いずれも標準的である。気象学を作らない──使うのは一つの指数関数と、その逆関数だけである。カオスを論じない──ローレンツ方程式も、アトラクタも、分岐も一切扱わない。数値予報を論じない──格子解像度も、パラメタリゼーションも、データ同化も扱わない。誤差が厳密に指数増大すると言わない──e^lambda t が成り立つのは誤差が小さいあいだだけであり、飽和に近づけば増大は止まる。本稿の計算は線形増大の領域に限る。 tau_d の値を主張しない──1.5 日は文献で広く用いられる代表値であり、季節・領域・変数によって 1 日から 2.5 日程度まで動く。第5節はこの依存の大きさそのものを示す。単一の指数だと言わない──実際の大気には尺度ごとに違う成長率があり、小さい渦ほど速く育つ。単一の tau_d は粗い近似である。予報の改善を否定しない──現に予報は延びてきた。本稿が言うのはその延び方が対数的であるということだけであり、改善が無意味だとは言わない。既刊との関係:論文253 は時間が一本なのが法則ではなく「予言できる」という要求だと示した──本稿はその要求が、どこまでなら満たせるかを数にする。論文195 は「安定」が六つの別の言葉だと分けた──あちらは安定性の分類、本稿は予測可能性の時間尺度であり、同じ双曲性を材料にして問いが違う。論文190 は「稀」を対数の目盛りで測った──本稿の見返りも対数である。論文266 は標本化定理の前提が決して満たされないと示した──本稿の「初期値を正確に知る」も決して満たされない前提であり、構図が同じである。加えたのは一日あたりの代償を 1.5874 倍という一定倍率として書いたこと、14->21->30->60 日の必要精度を 25.40/1625.5/1.70x10^9 倍と計算したこと、観測 10 倍が 4.98 日にしかならないと逆から書いたこと、tau_d を 1.0 から 2.5 まで振って答が 65536 倍から 84.4 倍まで動くと示したことである。 第一に、一日ごとの代償は一定倍率である。 tau_d=1.5 日なら、一日延ばすたびに初期値の精度が 1.5874 倍要る(第2節)。 第二に、これが本稿の芯である。14 日を 21 日にするのに 25.40 倍、30 日にするのに 1625.5 倍、60 日にするのに 1.70x10^9 倍(第2節)。 第三に、逆から見ると見返りは対数的である。観測を 10 倍精密にしても、延びるのは 4.98 日だけである(第3節)。 第四に、10 億倍にしても 44.85 日である(第3節)。 第五に、約二週間という数がここから出る。初期誤差が飽和の 10^-3 なら、予報可能な期間は 14.95 日(第4節)。 第六に、分離子は「指数か多項式か」である。 tau_d を 1.0 日にすると同じ延長に 65536 倍要り、2.5 日なら 84.4 倍で済む──すべてが一つの数に乗っている(第5節)。 予報の限界を決めているのは、方程式でも計算機でもなく、誤差の二重時間 tau_d という一つの数である。 tau_d=1.5 日なら、一日延ばすたびに初期値の精度が 1.5874 倍要る──どこで延ばしても倍率は同じだが、延長は足し算で、代償は掛け算なので、一週間で 25.4 倍、二週間で 645 倍、一か月半で 17 億倍になる。逆から見れば、観測を 10 倍精密にしても延びるのは 4.98 日であり、10 億倍にしても 44.85 日である。よく言われる「約二週間」も、この一行から出る──初期誤差が飽和の 10^-3 なら 14.95 日。分けるものは一つ──tau_d そのもの。1.0 日なら同じ延長に 65536 倍要り、2.5 日なら 84.4 倍で済む。776 倍の違いが、たった一つの数から生まれる。だから予報を延ばす仕事と、tau_d を測る仕事は、同じ重さを持っている。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

Executable Memory and World Coupling: Code as Cognitive Interface in a Self-Modifying Simulation

Large language model (LLM) agents have recently explored executable memory—compiling agent memory into code snippets that an external LLM interprets at inference time. We argue that this paradigm remains tied to a single architectural choice: the executor is an external model, the memory is a personal profile, and the code never participates in the agent's own memory economy. We present a cognitive simulation engine in which executable code is stored as unit-level memory entries and executed by a deterministic rule engine inside the simulation itself. A memory entry carrying an EXPR: prefix is a small program—an arithmetic expression over engine parameters and state variables—interpreted each generation; its result feeds directly into the unit's behavioral circuits. Code memory participates in the engine's memory economy: entries decay, are reinforced by hits, are evicted by capacity limits, and pass the same verification gates as any mechanism. Units acquire executable fragments by foraging, coupling energy gain with behavioral information transfer. Experiments show that (i) code memory measurably alters survival dynamics (extinction-count growth reduced by roughly 97% at threat 1.0); (ii) the survival benefit of code is stratified by strategy—decay reinforcement confers +15 generations at threat 1.5, healing reinforcement +10, while aggressive threat clearance confers no gain (clearing danger memories also clears the fear that drives defensive behavior); (iii) beyond a critical threat intensity (3.0) no code strategy confers benefit—a measured capability boundary; (iv) external trigger coupling: a unit's code can read an external trigger state (cognition) and, when the external signal is present, deterministically clear its own threat memories while writing an externally observable trace—with the external signal absent, the same code is inert, demonstrating that perception is a necessary component of the response; (v) cognitive code is acquired, not inherited: newly born units without the code fragment cannot perceive the external state, making cognition an evolvable individual trait; and (vi) when defensive and adversarial code coexist, an arms race emerges from primitive operations alone. We also report an unexpected semantics of negative-valued code, its diagnosis, and its redesign as a candidate inhibitory mechanism. The architecture points toward self-modifying systems in which memory, behavior, perception, and robustness converge on a single executable substrate.

Yizhang Hu · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.