CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
A unified framework that evaluates the capability of models to automate and augment another agent's performance, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Pattaraphon Kenny Wongchamcharoen, K. Gulati, Min Min Fong et al.
· 0 citations