Skip to content

Author

Jing Huang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition,...

Guang-Zhan Huang, Yong-Shuo Zhang, Bing-Tao Fu et al. · 0 citations

MobileDreamer: Generative Sketch World Model for GUI Agent

This paper proposes MobileDreamer, an efficient world-model-based lookahead framework to equip the GUI agents based on the future imagination provided by the world model, which consists of textual sketch world model and rollout imagination for GUI agent.

Yi-Lin Cao, Yufeng Zhong, Zhixiong Zeng et al. · 14 citations

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs and organizes benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Framework...

Xuanle Zhao, Qiushi Sun, Jingyu Xiao et al. · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.