The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They...
Jing-Ying Zeng, Zhen-Wei Dai, Jin-Ning Li et al.· 0 citations
The analysis reveals that while current LLMs struggle with efficient search in complex problems, incorporating systematic search strategies significantly enhances their problem-solving capabilities, highlighting the need for improving LLMs’ search abilities for real-world applications.
Min-hua Lin, Hui Liu, Xian-Feng Tang et al.· ACM Transactions on Knowledg...· 1 citation
Taste-Bench is built, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks, and it is shown that taste can be trained.
Wen-Bo Pan, Zhi-Chao Liu, Shu-Jie Liu et al.· 1 citation
OmniRouting is the first large-scale benchmark designed to evaluate LLMs on printed-circuit-board (PCB) routing reasoning under real-world industrial design-rule, manufacturability, and connectivity constraints, and reveals substantial limitations of current LMMs in PCB routing.
Tai-Ting Lu, Kai-Yuan Lin, Zi-Wei Dong et al.· 0 citations
OmniCAD is introduced, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery, and tool-augmented agentic reasoning.
Mingjia Wang, Tai-Ting Lu, Zi-Wei Dong et al.· 1 citation
OmniMech is introduced, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data, and experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of d...
Tai-Ting Lu, Run-Ze Liu, Zi-Wei Dong et al.· 0 citations
Retrospective Harness Optimization is introduced, a self-supervised method that optimizes the agent harness using only past trajectories and alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.
Wenbo Pan, Shujie Liu, Chin-Yew Lin et al.· arXiv.org· 8 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.