Skip to content
Preprint

ECLAIR: A Causally-Grounded AI Framework for Scientific Discovery in Empirical Software Engineering

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

This study presents the first causally grounded structured methodology for embedding LLMs within the scientific method in SE, designed around the epistemological demands of empirical SE research, establishing a basis for rigorous AI-assisted research.

Abstract

The scientific method has long guided empirical research in Software Engineering (SE), but the complexity of modern software systems often hinders its systematic application. This paper introduces ECLAIR, a causally grounded AI framework that integrates Large Language Models (LLMs) into every stage of the scientific process, from hypothesis generation to analysis and interpretation. ECLAIR treats LLMs as active scientific agents operating under the principles of causal inference, within a human-in-the-loop design that safeguards against the risks of unsound automated reasoning. We demonstrate the framework through a case study examining how prompt design influences code generation accuracy in two LLMs. Results show that, for both models, instruction-style, longer few-shot, and signature-augmented prompts yield small negative causal effects on accuracy, illustrating how causal reasoning provides a principled foundation for explaining why software phenomena occur. This study presents the first causally grounded structured methodology for embedding LLMs within the scientific method in SE, designed around the epistemological demands of empirical SE research, establishing a basis for rigorous AI-assisted research.

View source

Similar papers

#machine learning Review Jul 2026

CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

A framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee is presented, where a proof is checked by a program rather than read by a referee.

Jiyuan Tan, Vasilis Syrgkanis · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
Preprint Aug 2026

Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-

Used as a capable servant, generative AI has greatly accelerated intellectual work, yet it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. Preserving that agency calls for mechanisms that augment human metacognition during AI-assisted work. We therefore propose the Synthesis-Analysis Reciprocity Model, which views intellectual construction as a reciprocal interaction between two cognitive functions. Synthesis selects and combines the components of the artifact; Analysis evaluates them critically against objective indicators and constrains the Synthesis that follows. Grounded in this model, we present the Vibe Compiler, a research-logic compiler that helps researchers turn vague intuitions (Vibes) into coherent research logic. The system attempts to compile those intuitions against a paper ontology of 16 academic parameters. It treats compilation failures as signs that logical components are missing. Rather than fill those gaps autonomously, it returns reflective questions that prompt researchers to develop the missing reasoning themselves. We further characterize the origins of structural gaps along two orthogonal dimensions: cognitive function (Synthesis versus Analysis) and executing agent (human versus AI). The four resulting types of origin give a principled way to identify where breakdowns in intellectual construction arise. Crucially, our design implements the type in which the AI probes its own synthesized output, itself driven by the user's Vibes, and thereby stimulates human metacognition. This choice raises researchers from passive"Makers"of the output into"Managers"who critically direct and validate what the AI produces. In a prototype on NotebookLM and Gemini, AI behavior depended less on prompting than on the structure of the knowledge supplied. The framework spans a learner layer and a researcher layer.

R. Mizoguchi, Tomoki Aburatani, Kento Koike et al. · 0 citations
Review Jul 2026

The Case for Vibe Modeling: A Missing Step in AI-Based Trustworthy Software Development

A student survey study is presented that examines perceptions of LLM output understanding, validation effort, trust and the perceived usefulness of vibe modeling across several AI-assisted development scenarios to inform future studies for trustworthy and explainable AI-based software engineering via vibe modeling.

Shalini Chakraborty, M. Mittermaier, Judith Michael · 0 citations
Review Aug 2026

Loop Engineering: Building Blocks, Adoption, and Impact

An exploratory review of the emerging gray literature, which largely agrees on what a well-engineered loop contains: triggered agent runs bounded by machine-checkable stop conditions, persistent state files, verifier sub-agents, token budgets, and defined points of escalation to humans.

Jai Lal Lulla, Vahram Nersesyan, Seyedmoein Mohsenimofidi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.