A key challenge in bringing embodied intelligence into the real world is transferring capabilities from simulation to reality and enabling agents to continually adapt after deployment. End-to-end vision-language-action policies provide strong manipulation capabilities, but their transfer to physical environments typica...
Chen-Xi Li, Zhang-Rui Zhao, Rui Li et al.· 0 citations
Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exa...
Jun Yao, Ying-Fan Hua, Rui-Kun Li et al.· 0 citations
This work designs five types of multimodal tasks across text, molecular SMILES strings and images, and curates the datasets, demonstrating the feasibility of unifying multiple cross-modal chemical tasks within a single foundation model and enabling more intuitive, visual human-AI interaction.
Qian Tan, Di Zhang, Ben Gao et al.· arXiv.org· 16 citations
StraTA is a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL) and trains strategy generation and action execution jointly with a hierarchical GRPO-style rollout design, further enhanced by diverse strategy rollout and critical self-judgment.
Xiangyuan Xue, Yifan Zhou, Zidong Wang et al.· arXiv.org· 1 citation
An injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior, and a readout that factors cell energy into support salience and a competitive operation posterior are introduced.
Zhong-Yao Wang, Wan-Li Ouyang, Tao-Yong Cui et al.· 0 citations
Powder X-ray diffraction (PXRD) is the routine probe of crystalline matter, yet its analysis is the rate-limiting step as laboratories automate acquisition. Deep-learning analyzers excel on simulated patterns and degrade on measured ones. This simulation-to-real gap is structural, not additive: synthetic denoising give...
Shaoguang Wang, Weiyu Guo, Ben Fei et al.· 0 citations
World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic framework based on ortho...
Tao-Yong Cui, Zhong-Yao Wang, Xin-Yue Xu et al.· 0 citations
Immune therapies act across cell-intrinsic programs, tissue ecosystems, and patient-specific immune states, yet most predictors address these scales separately. We used a governed evolutionary AI Scientist to construct the Immune World Model, an action-conditioned model that learns how interventions move immune states...
Tao-Yong Cui, Xi Wang, Zong-Hang Li et al.· 0 citations
CENO is introduced, a family of long-context generative genomic world models designed to preserve local DNA grammar while extending usable context to regulatory and chromatin scales and provides a genome-scale sequence world-model framework for sequence interpretation, evolutionary reasoning, gene-scale reconstruction...
SciOrch is presented, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning, and attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.
Jingru Guo, Xiangyuan Xue, Lian Zhang et al.· arXiv.org· 0 citations
This survey systematizes recent progress in tree-search-based reasoning, viewing inference as instance-specific optimization rather than decoding, and introduces a Unified Design Space spanning search topology, evaluation signals, and control dynamics to unify a fragmented literature.
Jia-Qi Wei, Xiang Zhang, Yue-Jin Yang et al.· 0 citations
Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.