Skip to content
Conference

GenCTL: Constraint-Guided LLM Generation of Verifiable CTL Specifications from Natural Language Requirements

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · 0 citations

TL;DR

GenCTL is proposed, a prompt-based framework for effective and model-aware NL2CTL translation without task-specific fine-tuning that improves the reliability and practical checkability of LLM-generated CTL specifications.

Abstract

Natural language (NL) requirements are widely used in system design, but their ambiguity makes formal verification difficult. Translating NL requirements into Computation Tree Logic (CTL) is challenging because generated formulas must be both semantically appropriate and compatible with concrete verification models. Although large language models (LLMs) provide a promising basis for NL-to-CTL(NL2CTL) translation, their outputs are often unstable and prone to parsing, typing, and grounding errors. To address these issues, we propose GenCTL, a prompt-based framework for effective and model-aware NL2CTL translation without task-specific fine-tuning. GenCTL combines structured prompting, retrieval-enhanced few-shot examples, atomic-proposition (AP) grounding from nuXmv models, multi-candidate generation with frequency-first selection and length-normalized log-likelihood tie-breaking, and interactive refinement through an explanation dictionary. Experimental results demonstrate the effectiveness of the proposed framework in both model-agnostic and model-aware settings. On a generated NL--CTL dataset, the best automatically selected translation achieves 68% accuracy, which increases to 92% after one round of user refinement. On 150 NL requirements grounded in three nuXmv models, AP-list prompting yields 139/150 directly checkable formulas, increasing to 149/150 after lightweight normalization. These results show that GenCTL improves the reliability and practical checkability of LLM-generated CTL specifications.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parse...

J. Rosa, Pedro Santos, Valdemar Oliveira et al. · 0 citations
#software testing Preprint Sep 2026

Enhancing Automated Unit Test Generation for NLP Libraries Using Large Language Models

LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.

Amirhossein Deljouyi, Annibale Panichella, Andy Zaidman · 0 citations
Preprint Aug 2026

Automatic Translation of Unstructured Requirements into Linear Temporal Logic through Large Language Models

The results indicate that current general-purpose LLMs can achieve practically significant performance on the unstructured NL-to-LTL task without task-specific fine-tuning, and suggest that modern LLMs are becoming viable front-end assistants for semi-automated formalization workflows.

Alexandra Newcomb, Omar Ochoa · 0 citations
Preprint Aug 2026

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

ISO-grounded enrichment improves static quality proxies and reduces sensitivity to prompt wording, but does not reliably improve functional correctness; for error handling, extended-test pass rate decreases, suggesting tension between defensive coding patterns and exact-output benchmarks.

J. Pereira, V. Garcia · 0 citations
Review Open access 2026

When to Invoke an LLM in Industrial Natural-Language-to-DSL Conversion: A Taxonomy-Driven, Confidence-Gated Selective Hybrid for Safety-Critical Avionics Test-Script Authoring

Converting Korean natural-language avionics test procedures into an executable XML domain-specific language (DSL) is a labor-intensive bottleneck, yet neither automation extreme is acceptable in this safety-critical setting: end-to-end large language model (LLM) generation is unauditable and emits schema-violating valu...

Sungshin Lim, Hyuk-chul Kwon, Minho Kim · 1 citation
Preprint Aug 2026

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement.

Chit-Fung Lam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.