Skip to content
Preprint

VSpector: Specification-Driven Bug Detection for RISC-V CPUs

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

VSpector is presented, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns.

Abstract

Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts. In this paper, we leverage the official,natural-language RISC-V specifications as an effective information source for bug detection. We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns. To resolve the key technical trade-off between broad context scope and model reasoning accuracy when using Large Language Models (LLMs), VSpector employs a stepwise context refinement scheme across a four-stage pipeline: rule extraction, implementation localization, candidate identification, and sequential violation auditing. We evaluate VSpector on two industrial-strength RISC-V CPUs, CVA6 and XiangShan. Out of 217 reported candidates, manual inspection confirmed 148 true violations, representing a 68.2% precision. These violations correspond to 73 distinct bugs, including 42 previously unknown bugs. In our comparative experiments, DiveFuzz, a state-of-the-art CPU fuzzer, detected none of these new bugs during 24-hour runs per CPU. All 42 new bugs have been reported upstream, with developers already fixing 19 and confirming an additional 11 (30 in total), demonstrating that specification-driven auditing is a practical and complementary strategy for CPU bug detection.

View source

Similar papers

Open access Oct 2026

Towards Understanding the Bugs in Verilator, a Hardware Description Language Compiler

Verilator is the premier open-source Hardware Description Language (HDL) compiler. It transforms Verilog and SystemVerilog designs into optimized C++ or SystemC models, enabling high-speed, cycle-accurate simulation prior to large-scale production. As a cornerstone of the hardware verification ecosystem, the correctnes...

Song-Yan Jiang, Mao-Lin Sun, Kang Chen et al. · 0 citations
Preprint Sep 2026

C-to-Rust Fallacy: Automatic Refactoring != Memory Security

Rust has emerged as the leading system programming language, offering strong memory and type safety guarantees without compromising performance. This positions it as a compelling alternative to traditional languages like C and C++, which are susceptible to memory security bugs. However, manually transforming C to Rust...

Hung-Mao Chen, Xu He, Bo Lu et al. · 0 citations
Open access Aug 2026

Improving Bug Detection in LLM-Generated Unit Tests: Revisiting Test-Oracle Reliability Across Modern Large Language Models

This paper presents a formal mathematical model for categorizing the outcome of generated-tests into four classes, a couple of basic metrics: Bug-Revealing Rate (BRR) and Bug-Validating Rate (BVR); and two basic statistical tests to ensure that the results are rigorous.

Zeyad Farooq Lutfi · 0 citations
#software testing Book Open access Sep 2026

All Your Assembly Belongs to Rust: Automated Lifting for Uniform Testing and Verification

This paper translates Rust code containing RISC-V inline assembly into pure Rust code by emulating each instruction using a machine model extracted from the official RISC-V Sail ISA specification, and demonstrates how each category is handled by the translation.

Charly Castes, Gurvan Debaussart, Thomas Bourgeat · 0 citations
Preprint Aug 2026

InSPECtor: Improving SLEIGH Processor Specification Veracity via Proxy

This study is a significant undertaking to enable, for the first time, the systematic validation of open-source SLEIGH language specifications, predominantly used by Ghidra, with a testing framework based on an automated oracle validation strategy by proxy.

Michael Chesser, Paul Quirk, D. Cooke et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.