Skip to content

Revisiting TuRTLe: A Comprehensive Evaluation of LLMs for RTL Generation

Jul 2026 · ACM Transactions on Design Automation of Electronic Systems · 0 citations · 11 references

TL;DR

TuRTLe, a unified evaluation framework designed to systematically assess LLMs across key RTL generation tasks, is proposed, enabling a comprehensive assessment of LLM performance in syntax correctness, functional correctness, synthesis, PPA optimization, and exact line completion.

Abstract

Rapid advancements in LLMs have driven the adoption of generative AI in domains like Electronic Design Automation (EDA). Within the field of software development, EDA presents unique challenges derived from specific requirements of generated RTL code; RTL code must not only be syntactically correct and functionally accurate, but also synthesizable by hardware generators, while matching performance, power and area (PPA) constraints. These additional requirements introduce complexities that existing code-generation benchmarks often fail to capture, limiting their effectiveness in evaluating LLMs for RTL generation. To address this gap, we propose TuRTLe, a unified evaluation framework designed to systematically assess LLMs across key RTL generation tasks. TuRTLe integrates multiple existing benchmarks and automates the evaluation process, enabling a comprehensive assessment of LLM performance in syntax correctness, functional correctness, synthesis, PPA optimization, and exact line completion. Using this framework, a diverse set of forty open LLMs are assesed, tracking their strengths and weaknesses in EDA-specific tasks. Our results identify the best match for specific tasks (e.g., base models are better in module completion tasks, instruct-tuned models are better in specification-to-RTL tasks), while finding that recent models with autoregressive reasoning chain perform the best overall. We also analyze common compiler and runtime failures, study correlations between benchmarks and evaluation goals, and investigate potential training-data contamination in existing RTL datasets. These analyses provide further insight into the capabilities and limitations of current benchmarks for RTL generation.

View source

Similar papers

Jul 2026

Benchmarking LLMs for Verilog Design Flows

A reproducible benchmarking platform that evaluates open-source LLMs on Verilog RTL generation across 50 curated tasks consisting of combinational, sequential, finite state machine (FSM), and mixed designs, enabling reproducible evaluation of generative AI for hardware design workflows.

Angshuman Chakravertty, Rahul Koshti, Buddhi Prakash Sharma et al. · 0 citations
#small language model Preprint Sep 2026

Code Transformation Rule Synthesis using LLMs: Potential and Limits

Due to their black-box nature, LLMs suffer from limited explain- ability and a lack of determinism. Their usage cost can also rise, particularly with repetitive tasks on large codebases. To mitigate this, we conduct a novel empirical study targeting three domain- specific languages for transformation rules, namely Comb...

Axel Allain, Aymeric Blot, D. Khelladi et al. · 1 citation
Preprint Aug 2026

Route-Align-Verify for Functional Correctness in Code Generation

The results indicate that functional correctness in code generation can be meaningfully improved without modifying the backbone architecture, by jointly optimizing how tasks are prompted, how the model is adapted, and how final outputs are selected.

Eric Zhou, Jing Meng, Ao-Fan Liu · 0 citations
#machine learning Review Sep 2026

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is wri...

Fnu Aditi · 1 citation
#large language models Book Open access Oct 2026

From Unfinished to Done: Bridging the MDE Implementation Gap with Constrained LLMs

Model-driven engineering (MDE) excels at ensuring architectural consistency and managing complexity, yet extending generated code skeletons with project-specific business logic often remains a time-consuming manual task. Conversely, large language models (LLMs) offer immense flexibility in coding but may produce vibe-c...

J. C. Kirchhof, Lukas Netz, Dennis Mertens et al. · 0 citations

Automated Generation of RISC-V Extensions with Formal Correctness Guarantees

Janus is presented, an LLM-assisted framework that synthe-sizes custom instructions integrated into the Ibex RISC-V core while keeping correctness outside the agent, demonstrating a practical path for using LLMs to explore ISA specialization without making the agent part of the trusted correctness boundary.

Elisavet Lydia Alvanaki, Jia-Kun Wang, Eugenio Muscinelli et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.