Skip to content
#generative ai Open access

Supplementary Dataset and Benchmark Logs: From Semantic Retrieval to Conversational Agent

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This repository contains the supplementary materials and experimental data supporting the research article: "From Semantic Retrieval to Conversational Agent: A Web-Based RAG Architecture for Interactive System Dynamics Modeling". The dataset is divided into two primary components: the source model environment (search space) and the raw experimental benchmarks evaluating retrieval performance across different levels of user expertise and conversational search strategies. 1. Model Corpus, Queries, and Scenarios This section contains the definitions, domain classifications, and configurations used to build the semantic search environment and simulate user interactions. System Dynamics Models: Contains the extracted, curated, and serialized structural definitions of 63 System Dynamics models. These models cover diverse application domains, including Ecology, Macroeconomics, Smart Cities, Agriculture, and Epidemiology. User Queries and Intents: A dataset contrasting authentic broad novice search intents (e.g., "Show me health-related models") with theoretically perfect, expert-formulated structured queries requiring specific domain vocabulary. Benchmark Scenarios: 37 standardized benchmark scenarios engineered to evaluate cross-disciplinary semantic and lexical search performance across the system. Relevance judgments were established a priori by two domain experts, independently of any system output, and comprise 95 scenario–model relevance pairs 2. Experimental Benchmarks The benchmark execution logs provide a quantitative comparative analysis of different retrieval paradigms, running on a local AI ecosystem with direct CPU inference. File: conversational_rag_benchmark_metrics.csv: This file contains the aggregate Information Retrieval metrics (Precision@5, Recall@5, MRR, and nDCG@5) calculated for the 37 test scenarios. Note that the MRR_Mean column is computed over the full retrieval list (L = 10), whereas the paper reports MRR at the evaluation cutoff k = 5; two rows are affected (scenario 18, Method C, 1/7; scenario 37, Method E, 1/6), where the first relevant document falls beyond the top five, and setting both to zero reproduces the Table 5 values exactly. File: inference_latency_logs.csv: Documents the execution timestamps and hardware latency logs for the local ONNX inference engine, tracking the multi-turn conversational delays. File: ablation_study_p_values.csv: Contains the statistical hypothesis testing (paired t-tests) results validating the significance of the agentic retrieval improvements. File: contextless_retrieval_test.csv: Contains the isolated experimental data evaluating the impact of conversational memory (Method F). Evaluated Methodologies (Ablation Study) The benchmark data tests the following six retrieval paths: Method A: Broad Intent (Direct Retrieval Baseline) using standard single-turn semantic search. Method B: Agentic Refinement (Real Multi-Turn Agent Path) representing the complete conversational architecture. Method C: Expert Semantic Baseline (Direct Retrieval), establishing semantic search performance under optimal input conditions. Method D: Apache BM25 (Lexical over Expert Query) testing exact keyword matching. Method E: Expert Query via Agent (Single Agent Turn) to assess system robustness against over-complication. Method F: Contextless User Refinement (Direct Retrieval), submitting the user's raw Turn 2 answer directly to the vector database, thereby bypassing both the conversational history and the generative query rewriting step. Key Finding - Retrieval Accuracy: Replacing the static search baseline (Method A) with the Agentic Orchestrator (Method B) improves mean nDCG@5 from 0.1066 to 0.4422, a rise of over 300%. Expressed as retrieval success, Hit@5 rises from 0.1892 to 0.5946. Key Finding - Lexical vs. Semantic Dynamics: Under optimal conditions with expert queries, exact lexical matching (Method D) outperforms dense retrieval on every reported metric, achieving an MRR@5 of 0.8784 and an nDCG@5 of 0.8053 against 0.6856 and 0.5750 for semantic search (Method C). Key Finding - Computational Latency: The logs document the latency overhead of local CPU processing. A complete multi-turn exploratory session (Method B) averages 59.33 s (SD = 23.02 s), whereas structurally complete expert queries (Method E) execute in 49.81 s (SD = 9.84 s). Each scenario was executed as an independent cold-start process, so these values include ONNX session initialisation and constitute an empirical upper bound rather than steady-state deployment latency.

View source

Similar papers

#generative ai Open access Sep 2026

The socio-ecological costs of AI: Toward socially responsible and sustainable communication practices

The adoption of generative artificial intelligence among communication practitioners and researchers surged after the launch of ChatGPT in November 2022, urging practitioners to critically engage in exploring pathways for fostering socially responsible and environmentally sustainable AI practices.

Emma Christensen · 4 citations · ⚡1
#generative ai Open access Aug 2026

Baiyuan GEO Platform: A Whitepaper on Building a SaaS for Generative Engine Optimization

An engineering whitepaper documenting the construction of Baiyuan GEO Platform (2024–2026), a SaaS system for Generative Engine Optimization. The system helps brands be cited accurately and consistently across ChatGPT, Claude, Gemini, Perplexity, DeepSeek, and 15+ AI platforms. Coverage: seven-dimension AI citation-rate scoring algorithm, AI-Bot-friendly shadow document delivery (AXP) on customer-owned domains, Schema.org three-layer entity knowledge graph, closed-loop hallucination detection & auto-remediation, F12 three-layer structural optimizer (V1 rule-based + V3.1 dual-engine AutoGEO + E-GEO), rag-backend-v2 LLM hallucination hardening (six defense layers), and platform SSOT chain (brand_faq / page_type / alerts unification). v1.1.2 (this version): substantially expanded chapters 14, 15, 16 in both Traditional Chinese (zh-TW) and English (en) editions — added new sections covering early hand-tuning failure modes, bidirectional rollback design, placeholder guard trigger story, patch order causal chain analysis, cross-tenant cache privacy boundary, breadcrumb 404 ghost incident review (42 days, ~3000 ghost URLs), cross-microservice SSOT boundaries, and 5 engineering lessons (takeaways) per chapter — totaling ~13,000 additional words across 6 chapter files. Also adds LinkedIn launch announcement drafts (4 versions: zh-TW/en/ja personal + zh-TW company). v1.2.0 (this version): adds Part VI — three new chapters (Ch 17 cross-border China GEO with a Hong Kong edge node, UA routing, ICP-free central compliance and bidirectional AI visibility; Ch 18 AXP HTML Mirror-First semantic-HTML shadow documents; Ch 19 a five-layer cache-invalidation architecture for zero-touch propagation) in Traditional Chinese and English; backfills the Japanese edition to full parity (ja chapters 14–19 added); and expands Ch 13 (multimodal GEO) across all three languages with VideoObject GSC parity + origin backfill, a same-origin copyright filter, and sitemap image/video extensions. Languages: Traditional Chinese, English, and Japanese — all complete through chapter 19. License: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).

Vincent Lin · 3 citations
#generative ai Open access Aug 2026

Generic Identifiability and Directed Containment for Strongly Tree-Child Level-2 Networks under the Kimura Two-Parameter Model: The Principal Positive Domain and Strict Continuous Time

We classify regular full-dimensional stochastic containment among binary standard semi-directed strongly tree-child level-2 phylogenetic networks under the Kimura two-parameter (K2P) model. On the principal positive Fourier domain D₊ = {(s,g): 02s−1}, a directed containment germ exists if and only if the two labelled networks are isomorphic after independently redirecting ordinary three-cycle factors. The same condition is equivalent to a common full-dimensional regular germ; in particular, no proper one-sided containment occurs. It follows that the semi-directed topology is generically identifiable modulo ordinary triangle redirection, and that its structural triangle class is exactly reconstructible away from a proper algebraic exceptional set. The proof combines displayed-quartet inequalities and exact whole-map identities, an exact two-sector bridge-fibre theorem, physical marginal submersions, localization, and a bounded graph-to-algebra classification of cycle and theta factors. The bounded classification is computer-assisted: every directed primitive relation, rank exclusion, restoration parent, transport, and one-/two-port probe is represented by an exact certificate with independent replay and mutation evidence. The classification transfers to the strict continuous-time domain 0<s<1, s²<g<1. For every n≥3, two weakly but not strongly tree-child level-2 networks have continuous-time K2P images sharing a regular germ of dimension 4n−3, proving sharpness of strong tree-childness. This record is the complete v1.0.5-r1 priority and reproducibility package: the 26-page article, 24-page reader supplement, compile-complete five-file source archive, deterministic 495-member referee/verifier archive, external archive-qualification report, checksum sidecars, and dual-license notice. The manuscript source is v1.0.5; revision r1 repairs only an auxiliary probe-current semantic binding and changes neither the theorem, manuscript, PDFs, nor frozen classification. The clean verifier replay passed 41/41 layers, and the focused semantic mutation suite rejected 20/20 attacks. Exact source bindings: package tag k2p-same-referee-package-v1.0.5-r1; annotated tag object 6c9c89d38f4f4cdc9c328d8bb1237458c617136d; commit e2f6e32e6fe885e90c8e83a8c5b00785e663a4ae; referee archive SHA-256 4564cd1f8cd95f670a2e0d9619babaf3c343762cfd8ceeb190cd17df72802889. Article, supplement, and certificate data are licensed under CC BY 4.0; verifier and build code are licensed under MIT. No specific funding supported this work. The author declares no competing interests. Generative-AI assistance and its verification workflow are disclosed in the article. No mixed-sign K2P classification is claimed.r

Alec Kriebel · 2 citations
#generative ai Open access Aug 2026

The deflated-Welch statistic: a closed-form, guaranteed-level test for heteroscedastic one-way ANOVA. (T_BB, m01a): Methods, confidence-set-geometry development, derivations, novelty, and reporting standard): reproducibility deposit

The deflated-Welch statistic: a closed-form, guaranteed-level test for heteroscedastic one-way ANOVA William J. Dwyer, MD, MPH, FAAP — Department of Mathematics and Statistics, University of Massachusetts Lowell. ORCID 0009-0004-0855-7222. Concept DOI (always resolves to the latest version): 10.5281/zenodo.21908169. What this is The reproducibility deposit for the deflated-Welch statistic T_BB, a closed-form, guaranteed-level test for heteroscedastic one-way ANOVA (the Behrens–Fisher problem for k ≥ 3 groups). Welch's test becomes liberal under skew and unstable variance weights at small samples; T_BB = Q(s²)·exp(−R) keeps the ordinary group means and buys a guaranteed level by deflating the Welch quadratic by a Berger–Boos scale-inflation radius R. Three operating points are provided: a fixedcalibrated radius (κ_s), a design-adaptive near-guarantee radius (closed-form polygamma Cornish–Fisher with a finite-nkurtosis guard), and a fully proved smallest-eigenvalue radius R_eig (Gaussian, extended under bounded kurtosis). What the deposit contains Manuscript (author + anonymized) and a derivations supplement (DA1–DA13) plus a long-form derivations companion, covering: why Welch fails under skew in closed form; the Berger–Boos deflation and its exact worst-case radius; the polygamma-cumulant Cornish–Fisher radius with saddlepoint-exact normal backbone; the excess-kurtosis tail term with its finite-n upper-confidence guard; the imbalance correction; the fully proved smallest-eigenvalue radius (with the k-group multiplicity fix, free-β optimization, and the proved-under-bounded-kurtosis widening); and the k-sample Behrens–Fisher null distribution. Interactive demonstrator rerun_cochran/honest_anova.html — computes raw-mean Welch, the fixed / adaptive / proved T_BB radii, the estimand-changing transform routes, and the full routing receipt in the browser, reproducing the deposited Python. Its engine is extracted as a standalone Node module (m01A_anova_engine.js) and checked cell-by-cell against Python across an 84-design taxonomy (verify_anova_engine_taxonomy.py/.js, max |Δp| = 0.00000). Reproducibility scripts (rerun_cochran/, rerun/) — every reported number traces to a named, deterministically-seeded script (size/power/surface, the calibration and information-limit decompositions, the proved-radius verification, the imbalance calibration, the skew-router branch, and the figures). Real-data evidence — anova_flip_scan.py scans 2,783 public one-way layouts (254 datasets): guaranteed T_BBwithholds ~41% of Welch-significant calls, concentrated where the weight-instability screen fires, and never manufactures significance (Table 7 / Figure 15). Figures and the deterministic deposit builder (fixed timestamps → stable md5). All evaluation is simulation-based; the one empirical component is the public-dataset scan, which uses only openly distributed data. Code is released under the MIT License; text and figures under CC BY 4.0. Version history (consolidated changelog) Published version DOIs are marked ✅; the concept DOI above always resolves to the latest. Staged versions were rolled into the next published one unless noted. v1.0.77 ✅ 10.5281/zenodo.22167690 (2026-08-30): CSDA guide-for-authors conformance — abstract trimmed to 247 words (from 284), keywords cut to 7 (from 11), the withholding highlight shortened to ≤85 characters, and the arXiv PDF/source regenerated. No change to methods, results, figures, or code. v1.0.76 ✅ 10.5281/zenodo.22167536 (2026-08-30) — AI-disclosure heading aligned to Elsevier. The manuscript's declaration heading is now "Declaration of generative AI and AI-assisted technologies in the manuscript preparation process" (was "Use of generative AI"); the disclosure body is unchanged. Prepared alongside an Elsevier-compliant cover-letter variant and an EM suggested-reviewer sheet (both kept outside the deposit). docx/pdf rebuilt; deterministic md5 refreshed. v1.0.75 ✅ 10.5281/zenodo.22167304 (2026-08-30) — Submission-sharpening pass. Graphical abstract + Elsevier Highlights; figures and tables renumbered into reading order with per-table Source clauses; the validity–power frontier (Figure 8) now carries the proved R_eig operating point (100% validity, size-adjusted power 0.613, merge_tbb_proved_frontier.py); new Section 7 "Recovering power by design" + Table 8 (rc_anova_power_by_design.py); and a live required-n calculator in honest_anova.html (per-group and total n for 80% power, "power now @ total n"), with a numeric-heading CSS fix and the engine re-verified against Python at 0.00000. v1.0.74 ✅ 10.5281/zenodo.22165892 (2026-08-29) — Proved-under-bounded-kurtosis radius (DA12.6). The proved non-normal widening now keys on excess kurtosis, √(1 + κ̂·(n−1)/(2n)), from the exact Var(s²/σ²) = 2/(n−1) + κ/n, so symmetric heavy tails (Student-t) are covered where the old skew form √(1 + 0.75·skew²) under-covered; tbbProvedswitched to the kurtosis form across the demonstrator, engine, and Python truth (re-verified JS-vs-Python at 0.00000); new rc_anova_kurtosis_proof.py + deep-dive. v1.0.73 ✅ 10.5281/zenodo.22165709 (2026-08-29) — Reconstructed & verified demonstrator engine (standalone Node module + taxonomy verifier, max |Δp| = 0.00000 across 84 designs; Yuen zero-variance fix; T_BB-routed presets both directions); series-impact deep-dive (the corrected R_eig k-group multiplicity gap also reaches m03 and m01t). v1.0.72 (2026-08-29) — Title set to "The deflated-Welch statistic…"; corrected + optimized proved radius R_eig (β/k multiplicity fix + β-optimization, DA12); real-data Welch-vs-T_BB flip scan (2,783 layouts; Table 7 / Figure 15) + demonstrator imbalance-factor fix; long-form derivations companion. v1.0.71 / v1.0.70 (2026-08-21) — Zhang normal-reference comparator benchmarked on the efficiency frontier (valid on only 24% of designs, in the calibrated-liberal cluster); k = 2 adaptive-radius case-study fold (design-scaling vs shape-keying distinction). v1.0.69 ✅ 10.5281/zenodo.22035826 (2026-08-20) — HTML R1/R2 presentation pass + Figure 9 adaptive per-cluster label merge. v1.0.68 ✅ 10.5281/zenodo.22033737 (2026-08-20) — Companion consolidation into a single six-column Table 6; Figures 11–14 harmonized into one story. v1.0.67 / v1.0.65 / v1.0.60 (2026-08-19/20) — Guarded-reference naming-collision fix; the 40,000-replication expanded-frontier pin (Table 3 + Figure 8) with the symmetric-heteroscedastic skew-router branch; the mean-preserving lightened-R_eig do-not-use fallback. v1.0.59 ✅ 10.5281/zenodo.21995320 (2026-08-18) — Reporting standard + honest_anova.html demonstrator re-aligned to the current T_BB methods paper. v1.0.57 ✅ 10.5281/zenodo.21986847 (2026-08-17) — Reviewer-comprehension pass (multi-paragraph abstract, contributions list, trimmed captions); proved radius R_eig added as a Table 3 scorecard row; corner tail-index correction (N−k)/2 (low-order moments exist in every deployed design). v1.0.56–v1.0.49 (2026-08-16) — The k-sample Behrens–Fisher corner-distribution program: two-moment scaled-χ² corner reference, derived corner cumulants, the secular-eigenvalue law + closed CGF + power-law tail, consolidated into derivations DA13 with a prior-art/novelty audit. v1.0.48 ✅ 10.5281/zenodo.21963458 (2026-08-16) — The unifying λ(z) correction (a smooth instability-keyed deflation strength). v1.0.45 ✅ 10.5281/zenodo.21962965 (2026-08-16) — Atomic sparsity index + bootstrap-t edge hardening + shape-aware pooled standardized-residual bootstrap (SA-PSRB); multivariate transfer to m03. v1.0.44–v1.0.41 (2026-08-16) — Shape-moment re-injection order (skew is the sweet spot), validated and hardened pooled standardized-residual bootstrap, atomic weight-noise probes. v1.0.40 ✅ 10.5281/zenodo.21961667 (2026-08-16) — Log-domain weight-stabilization probe (negative for stabilization; clarifies the size-adjusted oracle ceiling); includes the oracle-power gap decomposition (≈92% conservatism, ≈8% estimation). v1.0.37 ✅ 10.5281/zenodo.21961327 (2026-08-16) — Residual-bootstrap qualification of the shoot-out + the first proved Gaussian smallest-eigenvalue radius R_eig (DA12, the p = 1 specialization of the m03 theorem). v1.0.36 (2026-08-15) — Figure 11 T_BB-region colour fix (amber, matching the routing figures). v1.0.27 ✅ 10.5281/zenodo.21908170 — Earlier published baseline of the deposit. Provenance: every number traces to a named, deterministically-seeded script listed in the manuscript Declarations; the demonstrator engine reproduces the deposited Python to max |Δp| = 0.00000 across the taxonomy verification. License. Code and scripts in the deposit are released under the MIT License; text and figures under CC BY 4.0. Reuse is permitted with attribution to the author and citation of the concept DOI above. How to cite. Dwyer, W. J. The deflated-Welch statistic: a closed-form, guaranteed-level test for heteroscedastic one-way ANOVA. Reproducibility deposit, Zenodo. https://doi.org/10.5281/zenodo.21908169

William Dwyer · 2 citations
#generative ai Open access Aug 2026

Ten-Year Panel of Japanese Municipal Finance from the Local Government Financial Settlement Survey

This R script (make_kessan10_csv.R) converts the Local Government Financial Settlement Survey (市町村別決算状況調), published by the Ministry of Internal Affairs and Communications on its annual pages of local government financial status survey materials, into machine-readable CSV. The source workbooks are print-oriented Excel files with multi-row merged headers, issued as four separate files per fiscal year (overview and expenditure, for cities and for towns and villages). The script consolidates them into long-format panels carrying fiscal year and municipality type as columns, and also writes one file per fiscal year. The output of a run over ten fiscal years (FY2015–FY2024) is deposited alongside it: all 1,741 municipalities, with 33 overview indicators and 94 expenditure items classified by purpose, giving panels of 17,410 rows each. Every municipality and every year is checked for internal consistency: the components of each expenditure category sum to that category's total, and the sum of all categories matches the total expenditure reported in the overview table. All checks passed for all ten years. Amounts are in thousands of yen, as published; blank cells are left blank rather than filled with zero. The column structure of the source data does not change over the period covered. One definitional change affects the adjusted ratio of current expenditure to current revenue: for FY2020 and FY2021 the special bonds issued for deferred tax collection are removed from current general revenue as well. Four changes of municipality occurred: Tomiya and Nakagawa became cities in FY2016 and FY2018 respectively, each receiving a new municipality code; Sasayama was renamed Tamba-Sasayama in FY2019, and Aogashima was renamed in FY2018 in the written form of its name only, both keeping their codes. The code was written with generative AI: Claude (Anthropic) was used to write and revise it. The author has verified the output and takes responsibility for the content. Version 1.1 corrects the reading of the census population change column in the overview table, where a small negative rate written with the triangle sign used in Japanese official statistics was left blank instead of being read as a number. 56 cells across the ten years were affected; no other value changed.

Yasutoshi Moteki · 1 citation

Related blog posts