Accurate depth estimation from endoscopic images is fundamental to surgical navigation, augmented reality, and robotic assistance, yet depth foundation models trained on natural images degrade severely on surgical endoscopy, and existing adaptation methods require surgical training data for each new camera and clinical...
Sheng-Yao Wang, Jun-Pei Xue, Xin Zhong et al.· IEEE Transactions on Medical...· 0 citations
Neural audio watermarks are increasingly deployed in commercial speech generation systems to make AI-generated speech traceable, yet their robustness has been studied mainly under conventional signal distortions. Since a watermark can be regarded as imperceptible noise added to the speech signal, a natural question is...
Xin Zhong, Sheng-Yao Wang, Ling-Feng Yao et al.· 0 citations
PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency that help generate spatially consistent FOA audio while mai...
Ling-Feng Yao, Chenpei Huang, Xing-Ke Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.