Skip to content
Preprint

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

benchmark provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning, and provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning.

Abstract

Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning process. Each problem pairs one or more circuit diagrams with a self-contained question, a typed or semantically specified answer, and a reference worked solution. An evidence-first construction pipeline aligns questions, figures, and solutions, while a reasoning-oriented taxonomy organizes problems by circuit type and dependency depth. Evaluation combines conservative typed scoring with identity-blinded multi-model semantic consensus, retaining every problem in the denominator. Across three commercial chatbot systems and six open-source multimodal large language models, the highest-scoring system reaches 84.8\% accuracy. However, performance consistently deteriorates on long-horizon problems, and qualitative analysis exposes persistent failures in topology-to-target binding, physical conventions, and late-stage output propagation. \benchmark{} provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning. Code are available at GitHub - CircuitReason/CircuitReason1K.

View source

Similar papers

Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solv...

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations
Review Aug 2026

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

SABRE is established as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark, and the results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Zi-Xuan Lan, Luzhe Sun, Matthew R. Walter et al. · 0 citations
Review Aug 2026

Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical understanding as a capability distinct from hig...

Ao-Xin Ni · 1 citation
Review Open access Jul 2026

Symbols and Neurons: A Review of Symbolic XAI in Deep Learning

A systematic review and synthesis of symbolic explainable AI (XAI) for deep learning is provided and a conceptual framework is proposed that clarifies training–inference flows, explanation interfaces, human feedback, and governance touchpoints is proposed.

Eduard Ionel Stan, G. Sciavicco, Paolo Napoletano · 0 citations
Jul 2026

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs'ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The test set entirely consists of composite problems: groups of single-domain subproblems that...

Liam Swayne · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.