This preprint investigates whether a fixed compound inference intervention can improve the performance of a small language model without updating its parameters. The experimental program compares structured prompt replacement, natural-language rewriting, source-preserving semantic augmentation, and a category-adaptive Qwen inference profile. Evaluation uses 48 IFBench and 48 LiveBench tasks with research-authored semantic-stress variants, programmatic scorers, pinned model and benchmark revisions, and matched stochastic seeds.In the prospectively specified P1.3 seed replication, Qwen3-1.7B improved from a mean objective score of 0.272 under raw non-thinking inference to 0.360 under source-preserving augmentation combined with category-adaptive inference. The paired effect was +0.088 (95% percentile-bootstrap CI: 0.018 to 0.160). Component analysis found a positive adaptive-inference-profile effect of +0.061, while the incremental contribution of semantic augmentation under matched adaptive inference remained uncertain at +0.027 (95% CI: −0.046 to 0.099).The study confirms the complete intervention on the fixed 96-task set under new stochastic draws; it does not establish independent task-level replication, cross-model generalization, or semantic augmentation as the active causal component. The package also increased prompt length, latency, and test-time computation. Version 1.1 provides detailed inference settings, benchmark identifiers, systems-cost accounting, causal boundaries, representative interventions, and reproducibility information.
Thibaud Peverelli· Zenodo (CERN European Organi...· 0 citations
This preprint investigates whether a fixed compound inference intervention can improve the performance of a small language model without updating its parameters. The experimental program compares structured prompt replacement, natural-language rewriting, source-preserving semantic augmentation, and a category-adaptive Qwen inference profile. Evaluation uses 48 IFBench and 48 LiveBench tasks with research-authored semantic-stress variants, programmatic scorers, pinned model and benchmark revisions, and matched stochastic seeds.In the prospectively specified P1.3 seed replication, Qwen3-1.7B improved from a mean objective score of 0.272 under raw non-thinking inference to 0.360 under source-preserving augmentation combined with category-adaptive inference. The paired effect was +0.088 (95% percentile-bootstrap CI: 0.018 to 0.160). Component analysis found a positive adaptive-inference-profile effect of +0.061, while the incremental contribution of semantic augmentation under matched adaptive inference remained uncertain at +0.027 (95% CI: −0.046 to 0.099).The study confirms the complete intervention on the fixed 96-task set under new stochastic draws; it does not establish independent task-level replication, cross-model generalization, or semantic augmentation as the active causal component. The package also increased prompt length, latency, and test-time computation. Version 1.1 provides detailed inference settings, benchmark identifiers, systems-cost accounting, causal boundaries, representative interventions, and reproducibility information.
Thibaud Peverelli· Zenodo (CERN European Organi...· 0 citations