Skip to content

Author

Meiyi Qiang

We have 3 of 19 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Hao Liang, Mingrui Chen, Hengyi Feng et al. · 0 citations
Preprint Jul 2026

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

WorkSurface-Bench is introduced, a benchmark for evaluating the capability of heterogeneous knowledge sources as surface routing in enterprise agents, and shows that correct surface selection is necessary but insufficient for task completion.

Hao Liang, Meiyi Qiang, Sizhe Qiu et al. · 0 citations
Jul 2026

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents across diverse scenarios with explicit state spaces, and introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis.

Chengyu Shen, Yujie Fu, Gang-Tao Xin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.