Skip to content
Preprint

AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

AQLoRA (Adaptive-Quantization LoRA), a recipe that reproduces Unsloth's hand-curated dynamic-4bit selection exactly, in seconds, where search-based allocation needs repeated calibration passes.

Abstract

Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of that time back. One CPU pass over the weights sets everything, with no search and no calibration data. The pass ranks layers by NF4 reconstruction error and keeps the top-K in fp16 under a memory budget. Those layers skip dequantization, which is where the speed comes from. A quality setting adapts every layer. A speed setting adapts only the top blocks, so the backward pass stops early. The rule reproduces Unsloth's hand-curated dynamic-4bit selection exactly, in seconds, where search-based allocation needs repeated calibration passes. We evaluate on Commonsense-170K across six models and four architecture families, from 1.4B to 14B. The speed setting trains 11.1 +/- 2.7% faster than well-tuned QLoRA and gives up about one accuracy point. It was faster in all nine independent timing sessions, at worst by 7%. The quality setting trains 4.8 +/- 2.4% faster. Its accuracy is level with QLoRA on every model and within a point of fp16 LoRA, for 0.2 GiB more memory. These error bars are measured between independent sessions, not within one. Earning them taught us three rules for timing on shared hardware. Fix the measurement duration, not the step count. Measure the noise floor from a duplicated arm, not a nearly identical method. Repeat whole sessions: a floor computed inside one sweep understates the real uncertainty several times over, and the random seed controls almost none of it. We validate the recipe with controls and report the two that failed. Choosing adapter layers by weight density is no better than random. Choosing protected layers by quantization error is not either. The count of protected layers, not their identity, carries the speed effect.

View source

Similar papers

Preprint Aug 2026

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.

Gunjun Lee, Sehwan Son, Younjoo Lee et al. · 0 citations
#machine learning Preprint Aug 2026

Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space

This work proposes code surrogate gradient as the first order signal in deployable code space to acceleate optimization, and performs guided search to preserve deployment faithfulness in fine-tuning low-bit models across different quantization datatypes.

Shiguang Wu, Zhou-Chen Lin, Quan-Ming Yao · 0 citations
#machine learning Preprint Sep 2026

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...

Jun-Hao Hu, S. Ramachandran · 1 citation
Preprint Aug 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.

Joao V. Cavalcanti, Ashia C. Wilson · 2 citations
Jul 2026

MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

It is shown that a layer's sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation, and proposed MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer, outperforms adaptive and mixed-precision baselines in ever...

Ashitabh Misra, Madhav Agrawal, Arham Jain et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.