Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full cor...
Soyeong Jeong, S. Jauhar, Sung Ju Hwang et al.· 0 citations
AI Night-Scientist is introduced, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning to suggest creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan et al.· 0 citations
This work critically examines the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges, and finds strong evidence that employing ensembles of diverse evaluation metrics consistently outperforms single-evaluator methods.
Anubhav Jangra, Bahareh Sarrafzadeh, Adrian de Wynter et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.