Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression
A three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model, with improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.
Hui-Cheng Zhang, Xi-Yao Feng, Ze-Tong Li et al.
· 0 citations