Skip to content

Author

Yangqiu Song

We have 8 of 59 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds

Large language model agents can invoke tools fluently, but enterprise workflows demand more than selecting the right tools: actions must strictly comply with organizational policies, tool feedback often conceals hidden side effects under partial observability, and long-horizon tasks require persistent state tracking ac...

Yi-Sen Gao, Yue (Sophie) Guo, Qing Zong et al. · 0 citations
#artificial intelligence Preprint Aug 2025

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. Although recent studies have developed frameworks to estimate PT parameters for Large Language Models (LLMs...

Rui Wang, Qi-Han Lin, Jiayu Liu et al. · 2 citations
#artificial intelligence Preprint Sep 2026

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

AgentIdeaBench is introduced, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration, and Scientific World Modeling is explored, a generation-time loop that refines a draft hypothesis through structured thought experiments.

Yunxiang Mo, Tianshi ZHENG, Yi-Sen Gao et al. · 0 citations
Preprint Aug 2026

Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V, and proposes PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition.

Bai-Xuan Xu, Yinyui Xu, Tianshi ZHENG et al. · 1 citation

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

Ka-po. Chung, Chunkit Chan, Yauwai Yim et al. · 0 citations
Conference Open access 2026

InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling

This work proposes InferenceDynamics, a flexible and scalable multi-dimensional routing framework by modeling the capability and knowledge of models, and demonstrates its effectiveness and generalizability in group-level routing using modern benchmarks including MMLU-Pro, GPQA, BigGen-Bench, and LiveBench.

Haochen Shi, Tianshi ZHENG, Weiqi Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.