Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in te...
Reading a base Qwen3-Omni with a logit lens at the audio-token positions, it is found that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token.
This work introduces Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol that improves direct CoT on multiple backbones and positions CFR as a targeted test-time intervention whose benefit depends on recoverable constraints and reliable Stage 1 extraction.
Hongbo Ma, Bang Yang, Y. Cheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.