Search plays a fundamental role in problem-solving across various domains, with most real-world decision-making problems being solvable through systematic search. Drawing inspiration from recent discussions on search and learning, we systematically explore the complementary relationship between search and Large Languag...
Min-hua Lin, Hui Liu, Xian-Feng Tang et al.· ACM Transactions on Knowledg...· 0 citations
OmniRouting is the first large-scale benchmark designed to evaluate LLMs on printed-circuit-board (PCB) routing reasoning under real-world industrial design-rule, manufacturability, and connectivity constraints, and reveals substantial limitations of current LMMs in PCB routing.
Taiting Lu, Kaiyuan Lin, Ziwei Dong et al.· 0 citations
OmniCAD is introduced, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery, and tool-augmented agentic reasoning.
Mingjia Wang, Tai-Ting Lu, Zi-Wei Dong et al.· 1 citation
OmniMech is introduced, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data, and experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of d...
Tai-Ting Lu, Run-Ze Liu, Zi-Wei Dong et al.· 0 citations
Retrospective Harness Optimization is introduced, a self-supervised method that optimizes the agent harness using only past trajectories and alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.
Wenbo Pan, Shujie Liu, Chin-Yew Lin et al.· arXiv.org· 8 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.