Skip to content

Author

Yulia Tsvetkov

We have 5 of 43 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

You're Hired: Strategic Model Selection for LLM Collaboration

While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systemat...

Zong-Wan Cao, Zi-Yuan Yang, Shang-Bin Feng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Multi-LLM Collaborative Alignment via Stackelberg Games

A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on whic...

Christina Hahn, Shang-Bin Feng, Dean Light et al. · 0 citations
#natural language process... Preprint Sep 2026

Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents

Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through...

Ji-Han Yao, Si-Han Zeng, Shang-Bin Feng et al. · 0 citations
#machine learning Preprint Sep 2026

Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL

Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, whe...

Zi-Yuan Yang, Yike Wang, Shang-Bin Feng et al. · 0 citations
Jul 2026

Rethinking the Evaluation of Harness Evolution for Agents

An extensive evaluation of automatic harness evolution for LLM agents is conducted, comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluating evolved harnesses on held-out tasks to assess whether the discovered improvements gen...

Yike Wang, Huaisheng Zhu, Zhengyu Hu et al. · 18 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.