Judge Arena: Benchmarking LLMs as Evaluators
Hugging Face Blog
· huggingface.co · November 19, 2024
Read on Hugging Face Blog →
Opens the original article in a new tab.
More from the blog
Hugging Face Blog
· huggingface.co
Aug 21, 2026
Measuring benchmark optimization in speech recognition
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Hugging Face Blog
· huggingface.co
Jun 30, 2026
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
MIT News · Artificial Intelligence
· news.mit.edu
Jun 26, 2026
LLMs help robots understand vague instructions and focus on key details
To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.