Skip to content

Author

Ao Qu

We have 5 of 35 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

MicrobeQuest: A Context-Aware Multimodal Benchmark for Microbiology Information Extraction

AI for Science (AI4S) is rapidly advancing scientific discovery, yet its progress critically depends on the availability of large-scale, high-quality structured scientific data and reliable evaluation benchmarks. In microbiology, abundant multimodal information—spanning text, images, tables, and charts—is distributed across scientific literature and public databases, making information extraction (IE) essential for transforming raw data into structured scientific knowledge. We formalize microbiology-specific IE as a class of tasks that require structured data extraction, joint multimodal understanding, and context-dependent scientific reasoning. Although numerous IE models have been proposed, their effectiveness in microbiological settings remains difficult to assess due to the lack of a standardized, domain-specific evaluation benchmark. To address this gap, we introduce MicrobeQuest, the first comprehensive multimodal benchmark for microbial information extraction tasks, comprising 11,877 expert-validated, context-aware query–response pairs. Evaluations of 19 state-of-the-art IE methods reveal substantial performance variability and modality-specific challenges, establishing MicrobeQuest as a standardized evaluation framework for advancing AI-driven microbiological research. All benchmark resources are publicly available at https://github.com/yulab-pku/MicrobeQuest.

Ou Zheng, Xue Ren, Xuexia Su et al. · 0 citations
Preprint Aug 2026

SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments as executable code with an OpenAI Gym-style reset()/step() interface, and a Reasoning Agent that learns to act in them. Each is a stateful, multi-turn environment (state transitions, reward functions, and verification code), so one interface spans reasoning problems and multi-step agentic tool use. The Reasoning Agent's regret is estimated using the gap between its reward with and without privileged hints; in optimizing this regret signal the Environment Designer learns to target environments at the edge of the agent's capabilities while keeping them feasible. Through extensive experimentation, we find several components critical to success: grounding the Environment Designer on documents sampled from a large pretraining corpus, and giving it an accumulated environment memory. Scaling to 30B-parameter models, SPADE improves over the strongest fixed-environment baseline by +5.3 on average across eight held-out math, science, code, and reasoning benchmarks, and lifts the tool-use setting by +5.7 on BFCL-v4 multi-turn and +13.9 on ACEBench-Agent; on the games setting, the margin over the strongest baseline grows with model scale. By making environment design itself a learnable component, SPADE takes a concrete step toward open-ended self-improvement.

Bo Liu, Simon Yu, Yiding Jiang et al. · 2 citations
#artificial intelligence Preprint Oct 2025

HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

The HUman-Grounded AGENT Benchmark is introduced, which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data.

Chance Jiajie Li, Zhenze Mo, Yuhan Tang et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

FAIRY is developed to execute and evaluate agentic agronomic operations on full-season spatiotemporal workflows that span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage, and an evaluation suite that combines agentic success, full-path spatiotemporal correctness, token cost, and edge-device runtime is developed.

Ao Qu, Panagiotis Michelakis, Lin-Yuan Han et al. · 0 citations
#artificial intelligence Preprint Aug 2026

SPADE: Self-Play in Adaptive Synthetic Executable Environments

SPADE (Self-Play in Adaptive Synthetic Executable Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments as executable code with an OpenAI Gym-style reset()/step() interface, and a Reasoning Agent that learns to act in them.

Bo Liu, Simon Yu, Yiding Jiang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.