Jul 2026· ACM Transactions on Software Engineering and Methodology· 0 citations· 43 references
TL;DR
ExReg, a human-in-the-loop workflow that automatically generates discriminative examples using SMT-based constraint solving, highlights how automated example generation guided by formal methods and mutations can improve the reliability, efficiency, and trustworthiness of LLM-assisted regex pattern generation.
Abstract
Regular expressions (regexes) are widely used in software development but remain difficult to author and validate due to their compact syntax and subtle semantics. While large language models (LLMs) can now generate regexes from natural-language descriptions, their outputs often miss developer intent, and existing refinement techniques rely on developers to craft positive and negative examples—especially discriminatory ones that expose fine-grained semantic differences. Producing such examples is cognitively demanding and often leaves regexes under-validated. This paper introduces ExReg, a human-in-the-loop workflow that shifts this burden away from developers. Given an ambiguous natural-language specification, an LLM first proposes multiple plausible regex candidates. Instead of requiring developers to devise discriminative examples, the system automatically generates them using SMT-based constraint solving. Developers need only affirm whether these synthesized test strings match their intent. The system further mutates candidates and reuses discriminative strings to systematically eliminate incorrect patterns. Across six benchmark datasets and six state-of-the-art LLMs, ExReg accurately identifies valid regexes—or determines that none are appropriate—while requiring only minimal developer validation. In particular, ExReg achieves an average accuracy of 87%, while requiring users to inspect only 5 examples on average, each with a mean length of 8 characters and an inter-example waiting time of under 13 seconds. These results highlight how automated example generation guided by formal methods and mutations can improve the reliability, efficiency, and trustworthiness of LLM-assisted regex pattern generation.
Most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.
This work formalizes Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels, and proposes and compares two modular verification strategies.
Tarfah Alrashed, Madhup Sukoon, David R Karger et al.· Proceedings of the VLDB Endo...· 0 citations
In fast-evolving software systems, effective 'natural language requirements parsing' and downstream change effect analysis capability across a multitude of codes represents low-hanging-fruit in this regard. We present a structured framework to deploy Large Language Models (LLMs) for automating two essential software en...
Nithya Krishnan, Kumaran Ramanujam, Suresh Babu Narra et al.· 2026 International Conferenc...· 0 citations
Experimental evaluation on 300 realistic pattern mining tasks demonstrates consistent improvements in algorithm configuration accuracy, parameter compliance, and dataset specification correctness across zero-shot, one-shot, and few-shot settings, highlighting the effectiveness of inference-time domain grounding for ena...
Madhavi Palla, Uday Kiran Rage, Arjun Chakravarthi Pogaku· International Journal of Dat...· 0 citations
This analysis covers 2,857 report-test-patch triplets from Defects4J and SWT-Bench using two widely adopted instruction-tuned LLMs from distinct model families and finds that both models exhibit systematic optimism relative to humans and only modest rank agreement, motivating bias-aware evaluation.
Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang et al.· arXiv.org· 0 citations
Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parse...
J. Rosa, Pedro Santos, Valdemar Oliveira et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.