Skip to content

Author

Dinesh Manocha

We have 6 of 16 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentan...

Sonal Kumar, Sinan Hersek, Artem Dementyev et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

This work develops ParA-LLM, a framework of 22 paralinguistic characteristics that surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench and releases ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories.

Nishit Anand, Jia-Qi Su, Ke Chen et al. · 1 citation
Preprint Aug 2026

DuplexWorld: Can voice agents help you get through the day?

D DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding, and shows that even the best voice agents leave substantial room for improvement on all 3 axes.

Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli et al. · 3 citations

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

It is shown that steering vectors can be sparsified by up to 85-96% while retaining most performance, and that different steering methodologies agree on a subset of important dimensions, and that different steering methodologies agree on a subset of important dimensions.

Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha · 6 citations
#artificial intelligence Preprint Aug 2026

VIBE: Video Instruction-aligned Background music gEneration

VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjec...

Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj et al. · 0 citations
#artificial intelligence Preprint Aug 2026

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the eval...

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.