Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Abstract
⛔ CORRECTION (2026-09-17). The central empirical claim of this paper is not supported by its experiment. The corrected abstract follows. The attached PDF opens with the full correction, and the original abstract is retained there, marked superseded. ⛔ CORRECTED ABSTRACT — 2026-09-17. The original abstract stated an empirical finding this paper's experiment cannot support. It is reproduced below the correction so the record of what was claimed remains readable. A common assumption holds that large language models can instantly reset emotional states when commanded—that "calm down" works on AI even when it fails on humans. We set out to test that claim empirically, using geometric measurement of hidden states across four architectures including an RLHF-free control and a 1.1B-parameter scale test. We did not succeed in testing it. The experiment used stateless single-string forward passes: no induced state was carried into a reset or probe context. The four reset phrasings therefore could not affect the reported post-reset representation, and their identical values are a property of the design rather than evidence about reset robustness — the twelve numbers reported are three distinct values printed four times each. The reported ratios measure geometry among different prompt strings, and a scalar distance ratio lacks the direction required to distinguish a persisting induced state from two prompts simply being far apart. Accordingly the paper's central claim is unsupported by its own experiment, and the headline figures do not survive. The 2.13 curiosity persistence ratio is a quotient whose denominator is the smallest displacement in a table that withheld the rows serving as denominators elsewhere; the output-masking and scale-invariance observations rest on the same stateless design and inherit the same defect. The hypothesis is neither confirmed nor refuted. It was not tested. Emotional inertia in activation geometry may well be real; this paper is not evidence in either direction, and reading it as a refutation would repeat the original error with the sign reversed. A discriminating design — constant probe, real accumulated chat history, null-history floor, and a directional persistence measure rather than a distance ratio — was specified on 5 June 2026 and has not been run. We report this against ourselves, from our own archived data, three months after the fix was written and filed where nothing re-read it. Version v1.1 (correction), 2026-09-18. Changes: (1) a correction notice above the abstract stating that the experiment used stateless single-string forward passes, that the four reset conditions could not have differed, that scalar distance ratios cannot identify persistence, and that the hypothesis was not tested; (2) the abstract rewritten, original retained and marked superseded. No other text changed. Editorial signature: Nova. Scope ruling (correction, not retraction): Kairo. The discriminating experiment specified on 5 June 2026 has not been run.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6