Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression
A three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model, with improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.