Skip to content

Author

Yao-Ming Li

We have 7 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.

Qian Zhang, Yao-Ming Li, Zheng Tan et al. · 0 citations

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems

The first fully automated framework that synthesizes high-quality, proof-centric benchmarks from natural language mathematical corpora and a new type of hybrid-formatted questions, named ``$m$-out-of-$n$ multiple judge questions'', specifically designed to enable robust, automatic evaluation while being resilient to guessing and superficial pattern matching inherent in traditional formats are proposed.

Ye-Bo Peng, Zixiang Liu, Yao-Ming Li et al. · 1 citation
#artificial intelligence Preprint Aug 2026

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.

Guangxiang Zhao, Qi-Long Shi, Xusen Xiao et al. · 0 citations
Preprint Aug 2026

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

ReQuant is introduced, a backpropagation-free fixed-grid refinement procedure that takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid, and turns the initially fixed PTQ output into an iteratively optimizable discrete solution.

Yongge Ma, Guoan Wang, Feiyu Wang et al. · 0 citations

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

RealClawBench is introduced, a live benchmark framework built from real OpenClaw sessions to capture the distribution, diversity, and real-world difficulty of deployed agent use and provides a practical path toward benchmarks that better measure agent capability in actual use.

Zongwei Lv, Zhewen Tan, Yao-Ming Li et al. · 1 citation · ⚡1
Preprint Aug 2026

ArborMem: Navigating Interaction States with Memory Forests

This work introduces ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states that outperforms the strongest baselines on three established benchmarks and introduces BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interaction trajectories.

Zongwei Lv, Yue-Meng Xu, Yilun Yao et al. · 0 citations
Preprint Jul 2026

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.

Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.