Synthetic Hospital is an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers and is a difficult and realistic test of clinical AI performance.
Christine Y. Park, Valerie Chen, Tim Dettmers· 0 citations
This paper draws on lessons from social-science accounts of human-human collaboration and argues that human-AI systems amplify these dynamics, introducing new asymmetries that make reasoning about uncertainty harder and introduce new coordination challenges.
Valerie Chen, Cleotilde Gonzalez, A. Woolley et al.· arXiv.org· 1 citation
DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, is introduced, and finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs.
Jenny T Liang, M. Bairathi, Wayne Chi et al.· arXiv.org· 1 citation
The results show that LLM agents, even without executing code, can identify many real-world reproducibility problems from paper-repository pairs, and ReproRepo can serve as a reusable, scalable framework for future evaluations of LLM agents on real-world reproducibility auditing.
Shanda Li, Qiuhong Anna Wei, Jingwu Tang et al.· arXiv.org· 0 citations
While agents aid initial task completion, they harm users' code comprehension and thus do not prepare users to extend their code, and low-effort agent interaction types, like copy+paste prompts and auto-accepted edits, are linked with lower comprehension.
VibeJam, a browser-based user study platform for users to collaborate with AI agents to develop websites, and open-source VibeJam to spur extensions and support studies on how coding agents can help users.
Nishant Balepur, Connor Baumler, Valerie Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.