Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventu...
Manato Yaguchi, Yotaro Kubo, Hikaru Asano et al.· 0 citations
This work proposes Feedback-to-Rubrics, a problem setting for learning criteria from inline comments on artifacts, which infers rubrics from these comments and iteratively refines them by observing errors in comment prediction based on the inferred rubrics.
Kotaro Yoshida, So Kuroki, Yuki Imajuku et al.· 0 citations
Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also s...
So Kuroki, Yujin Tang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.