LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect...
Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur et al.· 0 citations
This work presents SocialRL, a general recipe that trains social reasoning directly, and applies it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace, finding that in-domain training reaches the frontier.
Wenyue Hua, Zachary Huang, Tyler Payne et al.· 0 citations
This work presents Collaborative Reasoner, a framework to evaluate and improve the collaborative reasoning abilities of language models, and proposes a self-play method to generate synthetic multi-turn preference data and further train the language models to be better collaborators.
Ansong Ni, Ruta Desai, Yang Li et al.· Neural Information Processin...· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.