This work presents SocialRL, a general recipe that trains social reasoning directly, and applies it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace, finding that in-domain training reaches the frontier.
Wenyue Hua, Zachary Huang, Tyler Payne et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.