Skip to content

The Cheapest Digital Twin: A Pre-Registered Map of Twin-Native Extensions to a Frozen Single-Token Edge Encoder

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research) · 2 citations · 10 references

Abstract

The Cheapest Digital Twin: A Pre-Registered Map of Twin-Native Extensions to a Frozen Single-Token Edge Encoder Randolph James Ferlic, M.D. and Kimberly Kate Ferlic — Fieldstone Analytics, LLC, Austin, TX, USA Preprint · Zenodo DOI: 10.5281/zenodo.22922678 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract A frozen, deterministic, class-discriminant single-token encoder — features → a Fisher-discriminant ⊕ PCA subspace → a k-means codebook of ≤ 256 cells → a nearest-centroid decision emitted as one byte — has been characterized as a cost-versus-capability platform (a base token plus seven composable dials on one protocol). That platform framing left a prior question unanswered: the estate that produced the token described it, from the outset, as a deterministic, privacy-preserving functional twin of a signal source. If the byte-scale token is a digital twin, then the capabilities a digital twin is bought for — cheap transmission, fault-tolerance, anomaly detection against a personal baseline, a trust signal, anticipation, and low-cost refresh — should be reachable natively, without abandoning the frozen core. We test that hypothesis honestly. Six twin-native enhancements are run pre-registered, cross-domain on a shared seven-domain harness (ECG, industrial vibration, surgical-robot kinematics, wearable EMG, financial volatility, continuous glucose), IP-gated, and then hardened by a seven-test adversarial reviewer-proofing battery, a six-study strengthening round including a population-scale run over five ECG databases (133 patients), a twin-granularity sweep, and a self-mutation control. The result is a deliberately honest map. Three enhancements survive strong tests: a per-entity self-twin recovers a mean 92.8 % of a supervised per-patient anomaly ceiling without any labels, from as few as ten baseline beats, beating a shared population reference pooled (+0.076 macro-AUROC, 95 % CI [0.038, 0.117]) and on the majority of databases; an unsupervised alter-ego twin is a label-free error meter that turns a 76.6 %-accurate decision into ≈ 100 % accuracy at 70 % coverage where the decision's own confidence is uninformative, beating both a temperature-calibrated confidence and a supervised deep ensemble; and decision-gating — transmitting only when the decision changes — multiplies the token's already large data-movement advantage by a further 3.4–159× (decision-lossless), deepening rather than widening the moat. One enhancement is a clean negative (cross-channel imputation adds nothing over the filed omit-and-reweight decode), one is honestly bounded (anticipation is real only on slowly-evolving state — a glucose excursion is flagged 30–45 min ahead where a persistence baseline is worthless), and one is a boundary result (federated refresh helps under drift, not for same-distribution entities). A self-mutation ensemble — the tempting cheap substitute for the independent alter-ego — is shown not to replace it, isolating independence, not perturbation, as the trust meter's active ingredient. Every positive carries a pre-registered null or a stronger baseline that could have sunk it. The unifying thesis is commercial as much as scientific: the token is the cheapest digital twin — a byte-scale, deterministic, auditable, sub-milliwatt functional twin that delivers a useful subset of a digital twin's value at a fraction of a twin's cost. This is a characterization of previously described, filed methods; it discloses no new algorithmic subject matter, and the per-deployment selection of enhancements is retained as trade secret. Highlights · The reframe — the byte-scale token is the cheapest digital twin: a frozen, on-sensor, one-byte object serving a bounded, measured subset of the capabilities a digital twin is bought for, at microseconds and sub-milliwatt, with no raw signal on the wire. · The crown jewel — a label-free self-twin at population scale: from ten unlabeled baseline beats, a per-entity ≤ 64-cell twin recovers a mean 92.8 % of a supervised per-patient anomaly ceiling across five independent PhysioNet ECG databases (133 patients), beating a shared population reference pooled (+0.076 macro-AUROC, CI excludes 0) and on three of five databases — honestly database-heterogeneous, with per-entity granularity shown necessary. · A label-free trust meter that is actionable and beats the fair foils: an independent alter-ego's disagreement ranks the decision's errors where its entropy is dead (robot 0.500 → 0.944); as a selective-prediction policy it drives robot accuracy from 0.77 to ≈ 1.00 at 70 % coverage, beating calibrated confidence and a supervised five-model ensemble — and a self-mutation ensemble does not replace it. · The transmission twin deepens the data-movement moat: decision-lossless gating cuts transmissions 3.4–159× (with a shuffled-persistence null); composed with the one-byte payload advantage it reaches ≈ 1,000–3,000× on streams and ≈ 55,000× on steady monitoring — honestly a data-movement saving, not a total-energy one. · Honest negatives and boundaries are the credibility: imputation buys nothing; same-distribution refresh is flat; anticipation is real only on slowly-evolving state; the self-mutation trust meter fails. A seven-test reviewer-proofing battery strengthened three claims, narrowed two, and deflated one. · All characterization of filed/published methods — no new algorithmic subject matter; the per-deployment twin selection is a trade secret. What this record contains · `Manuscript_Paper48.pdf` — the manuscript, seven figures embedded (self-twin forest; operating envelope; granularity sweep; alter-ego selective prediction; glucose lead-time; energy composition; honest-map scorecard); and `Manuscript_Paper48.docx`, the editable source. · `PAPER_48_ZENODO_ARCHIVE.zip` — the reproducibility archive (md5 in ARCHIVE_MD5.txt): the eight frozen pre-registrations, the shared seven-domain harness and all runners (the three single-enhancement runners and their cross-domain matrix versions, the alter-ego and anticipation/refresh matrices, the reviewer-proofing battery, the EKG deep-dive + clean-INCART re-run + the population-scale/granularity Modal runner, the selective-prediction, gating-scaling- law, CGM lead-time, and mutant-ensemble runners, plus the energy ledger and figure builder), the two shared library modules, the per-experiment result records, the run-logs, the seven figures, the manuscript source, and a README. All datasets are public and not redistributed (fetched from their public sources at run time); all paths and identifiers are scrubbed (absolute paths → PATH_TO_DATA, cloud handle → MODAL_USER) and leak-scanned. Cite as R. J. Ferlic and K. K. Ferlic, "The cheapest digital twin: a pre-registered map of twin-native extensions to a frozen single-token edge encoder," Zenodo, 2026, doi: 10.5281/zenodo.22922678. License and patent notice Released under CC-BY 4.0. Consistent with that license, no patent or IP right of the authors is licensed, waived, or conveyed by this deposit. This work characterizes previously-described methods and discloses no new algorithmic subject matter; k-means / vector quantization, Fisher discriminant analysis, deep ensembles, bagging, query-by-committee, and nearest-centroid novelty detection are established prior art, used only as tools. The methods characterized are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), the personalization / on-device-adaptation applications (priority U.S. Application No. 19/467,303 and its continuations, incl. the ephemeral-cache application No. 64/142,667), and the transition-monitor / entropy-coding applications (Nos. 64/096,355 and 64/096,927). The per-deployment selection of which twin to dial in is retained as a trade secret and is not disclosed here. © 2026 Fieldstone Analytics, LLC and the authors. Inquiries: randolphf@fieldstoneanalyticsllc.com. Companion deposits (spiral-domain-encoder-campaign) · The Configurable Bottleneck (the platform this extends): doi:10.5281/zenodo.22884023 · Class-discriminant codebook construction (the base encoder): doi:10.5281/zenodo.20788187 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · Label-free inference-time channel fusion: doi:10.5281/zenodo.22046713 · The predictive reach of a decision token: doi:10.5281/zenodo.22736921 · Non-invertible but not anonymous (privacy): doi:10.5281/zenodo.22819210 · Unlinkable but not anonymous (privacy): doi:10.5281/zenodo.22838120 · The price of the bottleneck (deployment): doi:10.5281/zenodo.22838118 · Paying down the price of the bottleneck: doi:10.5281/zenodo.22866039 · Token as a generative-edge cache key: doi:10.5281/zenodo.22148612 Keywords digital twin; functional twin; edge AI; edge computing; TinyML; microcontroller; sub-milliwatt inference; decision token; vector quantization; k-means codebook; nearest-centroid classifier; information bottleneck; model compression; energy efficiency; inference latency; data movement; transmission energy; duty cycling; radio energy; wireless sensor networks; anomaly detection; novelty detection; out-of-distribution detection; personalization; patient-specific ECG; arrhythmia detection; heartbeat classification; label-free; unsupervised learning; few-shot learning; uncertainty quantification; calibration; conformal prediction; selective prediction; abstention; deep ensembles; bagging; query by committee; bootstrap confidence intervals; electrocardiogram; bearing fault diagnosis; surgical robotics; electromyography; continuous glucose monitoring; financial volatility; cross-database generalization; concept drift; fault tolerance; interpretable models; auditability; pre-registration; honest negatives; reproducibility

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.