Skip to content
Preprint

Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

Handoff-H1 is presented, a takeoff system built from three layers: purpose-built computer-vision models that extract primitives; tool-using agents equipped with image operations and in-house visual-task tools, including CV-model-backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base.

Abstract

Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state. We present Handoff-H1, a takeoff system built from three layers: purpose-built computer-vision models that extract primitives; tool-using agents equipped with image operations and in-house visual-task tools, including CV-model-backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base. We evaluate on the Construction Blueprint Takeoff Benchmark: 10 real residential blueprint sets paired with consensus-validated expert takeoffs - 2,009 verified line items, restricted for scoring to the 1,348 primary-tier materials that drive an estimate - scored per trade by an LLM judge on material coverage and quantity Precision@25% (P@.25) and combined into a weighted composite. Under identical scoring from the raw PDF, seven frontier and open-weight models span composites of 35-61, and independent professional estimators - scored against the same reconciled gold standard - post 77.6% (65.5% coverage, 87.9% P@.25). Handoff-H1, working end-to-end from the raw PDF, reaches 81.6% (86.1% coverage, 78.8% P@.25): roughly 20 points above the strongest frontier agent, and above the independent estimators by pairing near-human quantity precision with coverage they do not reach. The evaluation harness is public for the open harbor framework; the blueprint sets and ground truth are available upon request for research use.

View source

Similar papers

Preprint Aug 2026

Lift, Associate, and Fuse: A Decision-Centric Framework for 2D-to-3D Foundation Model Transfer

Methods that transfer predictions from two-dimensional foundation models into three-dimensional segmentation are commonly grouped by task or representation. Those groupings obscure the decisions that determine whether a system remains coherent across views: where image evidence is grounded, when observations become one...

Wen-Tao Sun, Yi-Ping Chen, J. Zelek et al. · 0 citations
Open access Sep 2026

A Knowledge-Driven Intelligent Agent for Automated Quantity Checking of Concrete Bridge Structures

Automated quantity checking directly from two-dimensional bridge drawings remains challenging because the required information is distributed across structural views, detail drawings, and tables, while recognition errors may propagate into deterministic engineering calculations. This study proposes a knowledge-driven i...

Yi Li, Bo-Xu Tian, Qing Liu et al. · 0 citations
#artificial intelligence Review Sep 2026

BlueprintAgent: Constraint-Triggered Targeted Revisits for Simulation-Ready Generation from Scanned Structural Blueprints

Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured frame representations that support deterministic FEM export and qualified-engineer review---underpins safety assessment and seismic retrofit, but the process remains manual. Direct prompting of a multimodal larg...

Zhou-Yuan Xu, Chen Yang, Lin-Hao Wang et al. · 0 citations
Preprint Sep 2026

WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery

Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A...

Zhong-Yu Pang · 0 citations
Open access Aug 2026

Vision-Based Digital Twin and AI Agent Framework for Low-Cost, Explainable Indoor Building Inspection and Safety Assessment

Aging residential buildings constructed under outdated design standards create an urgent need for scalable, evidence-based indoor safety assessment methods. Conventional manual inspections rely on subjective checklists, lack audit trails, and are impractical for widespread deployment. This study presents a vision-based...

Zijian Jing, Li-Ying Zhu, Tianyi Chen et al. · 0 citations
Preprint Aug 2026

SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling

Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy training and evaluation. Natural language descriptions are easy to read, but their linguistic variability makes automatic verification difficult. Rigid template formats, such as BEHAVIOR-1K's skill_annotation, a...

Yupu Lu, Shuang Wu, Si-Han Chen et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.