Pretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6B
Almost every language model in service was trained by one family of software. That concentration makes a question hard to settle: how much of what is known about training a language model describes language models, and how much describes that software? Settling it needs a second implementation able to carry a model through a whole lifecycle rather than reproduce one operator. We report such a lifecycle. Using numbat, a machine-learning stack written in Zig with no third-party runtime dependencies, we pretrain a 124.4 M-parameter GPT-2 from random initialisation over 9.91 B tokens of web text, then adapt a separate small model to clinical question answering. A reference implementation runs on identical hardware at both stages, and a sidecar with authority to halt a run supervises each. Agreement is close. Held-out cross-entropy finishes at 3.2588 against a published 3.29, and HellaSwag at 0.3053 against 0.299; across 8 paired evaluations it sits below a same-machine reference at every point, by 0.0608 on average. Re-running that reference on different hardware moves it 0.0035, which bounds how much of any gap is method rather than framework. Throughput does not pay for agreement: measured in one session at a production configuration, numbat reaches 43,374 tokens per second against PyTorch's 41,202, scaling 2.769x over three cards. Clinical adaptation ends at 2.1899 held-out loss against 2.1941. Neither model is a medical device, and neither is validated for clinical use.
: Small Language Models (SLMs) run faster and fit on modest hardware, yet solving multi-step logic problems has traditionally been difficult for them. This work investigates a systematic, multi-model framework that combines Chain-of-Thought (CoT) prompting with Low-Rank Adaptation (LoRA) parameter-efficient fine-tuning...
Aryan Raina, Shiwani Gupta, Jagruti Jadhav et al.· Journal of Computer Science· 0 citations
Whether independently implemented training stacks can serve as differential oracles for a whole fine-tuning pipeline, rather than the operators and inference paths that prior differential testing targets, is studied.
. Different large language models vary dramatically in their resource requirements, with some costing millions to train while others can run on a laptop. This paper examines four models that represent different design philosophies: GPT-3, GPT-4, LLaMA 2, and PaLM 2. GPT-3, with its 175 billion parameters, first showed...
Guo-Yu Xu· Proceedings of the 4th Inter...· 0 citations
Manac-a-1B is the strongest model below the 7B scale, exceeding both Tucano-1b1 and Tucano-2b4 on LAMBADA-PT with large paired margins and is competitive on commonsense completion and near chance on multiple-choice reasoning, as are all small base models.
Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fábio Porto· 0 citations
This work reports on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies, and treats a widely used reference implementation as an executable specification and verify against it at five levels.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.