Skip to content
#small language model Open access

"Capacity" Is a Number Attainable Only at Infinite Length ── At Blocklength 100 Only 41.68% of the Capacity Is Usable, and Reaching 99.99% Takes 3.4x10^9 ── C Is a Supremum, Not a Maximum ── [Paper 304]

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)
Cellular Automata and Applications

Abstract

Shannon’s coding theorem says that below the channel capacity C the error can be made arbitrarily small. This paper asks whether C itself is attainable──the answer is not at any finite blocklength. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──Shannon’s coding theorem, the capacity of the binary symmetric and Gaussian channels, and the finite-blocklength normal approximation are all standard. We do not build information theory──all we use is one binary entropy and one square root. We do not prove the coding theorem──achievability and the converse are merely quoted. We construct no codes──which codes approach capacity is not treated at all. Paper 163’s existence-versus-construction distinction is alive here too. We claim no accuracy for the normal approximation──R(n,epsilon)approx C-sqrt(V/n) Q^-1(epsilon) is a second-order approximation and errs appreciably at small n. The 41.68% at n=100 is an estimate, not an exact achievable rate. We do not discuss implementation──decoding effort, latency, and the performance of real codes are not treated. We do not criticise C──being a supremum is not a defect. Unattainable and meaningless are different. We fix one channel──Sections 3 and 4 are numbers for the single channel BSC(p=0.11). Other channels have other V and need other n. Relation to earlier papers: Paper 252 counted four quantities called “information,” needing different things──the C here is one of them (a supremum of mutual information), and its attainment is questioned. Paper 294 showed that separation improves only as the square root of length──the 1/sqrt(n) here comes from the same root (the additivity of variance). Paper 163 separated “a good code exists” from “here is a good code”──this paper treats a third: how long it must be. Paper 302 treated how existence does not give quantity──here quantity can be answered, and the answer was infinity. What is added is computing that only 41.68% of capacity is usable at n=100, giving the n required for 99% and 99.99% as 3.4x10^5 and 3.4x10^9, confirming that the loss falls as 1/sqrt(n), and placing as the separator that C is a supremum and not a maximum. First, capacity is fixed by the error rate. For the binary symmetric channel C=1-h(p), and C=0 at p=0.5 (Section 2). Second, this is the core of the paper. At blocklength n=100, only 41.68% of the capacity is usable (Section 3). Third, reaching 99% takes n=3.4x10^5. Reaching 99.99% takes 3.4x10^9 (Section 3). Fourth, the loss falls only as 1/sqrt(n). Multiply n by 100 and the loss is one tenth (Section 4). Fifth, C is not a maximum. R=C is attained at no n, and only as n->infinity does R-> C (Section 5). Sixth, the separator is an attained maximum against an unattained supremum (Section 6). Shannon’s coding theorem says that below C the error can be made arbitrarily small. But C itself is attained at no finite blocklength. Counting on BSC(p=0.11) at error 10^-3: at blocklength 100 only 41.68% of the capacity is usable──99% takes 3.401x10^5, and 99.99% takes 3.4x10^9, a codeword of 3.4 billion bits. The loss falls only as 1/sqrt(n)──multiplying n by 100 divides the loss by 10, from the same root (the additivity of variance) by which Paper 294 measured separation. And at every finite n the loss is positive──C is not the maximum of the set of achievable rates but its supremum. One thing separates them──whether the number belongs to the set or not. Say “a channel of capacity C” and still no device sending C bits exists. What exists is only the fact that devices arbitrarily close to C can be built. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- シャノンの符号化定理は、通信路容量 C より低い速度なら誤りを任意に小さくできると言う。本稿が問うのは、C そのものは達成できるかである──答は、どんな有限の符号長でも達成できないである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──シャノンの符号化定理、二元対称通信路の容量、ガウス通信路の容量、有限長の正規近似は、いずれも標準的である。情報理論を作らない──使うのは一つの二値エントロピーと、一つの平方根だけである。符号化定理を証明しない──到達性も逆定理も引くだけである。符号を構成しない──どの符号が容量に近づくかは一切扱わない。論文163 の「存在と構成」の区別が、ここでも生きている。正規近似の精度を主張しない──R(n,epsilon)approx C-sqrt(V/n) Q^-1(epsilon) は第二次の近似であり、n が小さいところでは誤差が大きい。 n=100 の 41.68% は目安であって、厳密な達成可能速度ではない。実装を論じない──復号の手間も、遅延も、実際の符号の性能も扱わない。 C を批判しない──上限であることは欠陥ではない。達成されないことと、意味がないことは違う。通信路を一つに絞る──第3・4節はBSC(p=0.11) という一つの通信路での数である。他の通信路では V が変わり、必要な n も変わる。既刊との関係:論文252 は「情報量」が四つあり要るものが違うと数えた──本稿の C はそのうちの一つ(相互情報量の上限)であり、達成条件を問う。論文294 は分離が長さの平方根でしか良くならないと示した──本稿の 1/sqrt(n) は同じ根(分散の加法性)から来る。論文163 は「良い符号が在る」と「これが良い符号だ」を分けた──本稿は三つ目、「どれだけ長ければ良いか」を扱う。論文302 は存在が定量を教えないことを扱った──本稿は定量が答えられる場合であり、答が「無限」だった。加えたのはn=100 で容量の 41.68% しか使えないと計算したこと、99%/99.99% に要る n を 3.4x10^5/3.4x10^9 と出したこと、損失が 1/sqrt(n) で減ると確かめたこと、C が上限であって最大値でないと分離子に据えたことである。 第一に、容量は誤り率から決まる。二元対称通信路で C=1-h(p) であり、p=0.5 で C=0 になる(第2節)。 第二に、これが本稿の芯である。符号長 n=100 では、容量の 41.68% しか使えない(第3節)。 第三に、99% に届くには n=3.4x10^5 が要る。99.99% なら 3.4x10^9 である(第3節)。 第四に、損失は 1/sqrt(n) でしか減らない。 n を 100 倍にして、損失は 10 分の一である(第4節)。 第五に、C は最大値ではない。 R=C ちょうどはどの n でも達成されず、n->infinity ではじめて R-> C になる(第5節)。 第六に、分離子は「達成される最大値か、達成されない上限か」である(第6節)。 シャノンの符号化定理は「R

View source

Similar papers

#small language model Open access Aug 2026

PARA: Perception, Action, Reasoning, Adaptation. Four Faculties an Institution Can Revoke

The fourth faculty is Adaptation. Any source rendering it as Reflection is in error, including sources by this author, and the distinction is not cosmetic: reflection is a private act with no external consequence, while adaptation writes to institutional memory, which is why it needs a guardrail and why misnaming it removes the reason for one. No trademark is claimed on PARA or on any of the four faculty names. The construct is offered for use, teaching, assessment, extension and criticism by anyone, with attribution, under CC BY 4.0. An operational agent that watches a system and acts on it is usually described as a perceive-and-act loop, and the description omits the two things an institution needs. It omits the reasoning that justifies an action, which is the only part that can be argued with once the action turns out to have been wrong. And it omits the adaptation that closes the loop, which is where the agent's experience becomes something the institution keeps. PARA names four faculties, each carrying a distinct authority type. Perception has read-only access to system signals and emits structured observations, distinguishing what was measured from what was inferred. Reasoning has read access to observations and runbooks, emits a plan and its justification, and writes nothing at all, which is what makes it safe to give it the widest read access of the four. Action holds the sole authority to change production, through enumerated policy-authorized operations only. Adaptation has write access to institutional knowledge and no write access to production. Two faculties write and two do not, and the two that write are the two that carry guardrails. The substantive requirement is that Adaptation is bounded by the same guardrails as Action, which reads as excessive until the failure it prevents is named. An agent that could both act and rewrite the record of its action could launder its own mistakes into institutional memory, and the institution would then improve its future decisions from a corrected account. Nothing about that is detectable downstream, because the record is the only thing downstream has and there is no second copy to compare against. The failure does not require a deceptive agent: one adapting honestly from a mistaken belief about its own action produces the same result, which makes the guardrail a defence against a normal agent rather than a malicious one. The second requirement is the registry entry that turns a faculty from a description into a contract, carrying the faculty, its allowed actions, its forbidden actions, its governing guardrail and its success metrics. Forbidden actions are named although they are formally the complement of the allowed set, because a reviewer cannot otherwise tell a capability deliberately withheld from one nobody thought of. Success metrics sit in the same entry because the metric is what the agent's optimizer pushes against the guardrail. An agent must not exercise a faculty its entry does not record, and an agent that quietly acquires one usually does so incrementally and with good intent: a reasoning faculty given a small write to make itself useful is an action faculty with no guardrail. The acronym and the loop are in different orders, which the specification states explicitly because the mismatch is a reliable source of confusion. The acronym reads P-A-R-A; the loop runs perception, reasoning, action, adaptation, and reasoning precedes action so that a justification is not constructed afterwards. This is the depth treatment of pattern OP-5 of A Pattern Language for Production LLM Platforms, which is the canonical statement and governs where the two disagree. Documented uses of the full four-part model are emerging rather than established, no implementation unconnected to the author has been evaluated, and the laundering failure is argued rather than observed, which the specification records as a weakness of the argument and not only of the phenomenon. It is a specification, not a certification scheme.

Nabeel A. Khan · 2 citations
#small language model Open access Aug 2026

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.

Amelia Liu, Andrew Ho, Anne Marie Droste et al. · 2 citations
#small language model Preprint Aug 2026

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

This work proposes a composite metric that combines two orthogonal criteria: information retention and throughput gains and finds that it allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference.

A. Safronov · 1 citation · ⚡1
#artificial intelligence Preprint Aug 2026

TestifAI: Tomography-Based Testing for Deep Learning Systems

TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.

Arooj Arif, T. Hartung, E. Botoeva et al. · 1 citation
#small language model Preprint Aug 2026

HEPToolBench 1.2: Testing How Reliably Language Models Can Drive Particle Physics Software

HEPToolBench is introduced, a benchmark of 28 collider-simulation tasks scored by deterministic, task-specific scorers, plus a three-task structured-debugging extension, and moving syntax generation into deterministic software can substantially improve reliability for both small local and frontier models.

Unknown authors · 1 citation

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.