Jul 2026· Electronics· Vol 15, pp. 3041· 0 citations
TL;DR
This paper describes how the CustomNerd system can be used to build an expertise-driven application through a running example, explains the system’s architecture, discusses related work, and evaluates CloudNerd on a 120-question Stack Overflow benchmark against OpenDeepResearch and AutoGPT using semantic textual similarity and RAGAS metrics with paired permutation tests.
Abstract
CustomNerd is a modular framework to build domain-specific expertise systems using a common set of core components. The system is designed to support multiple domains and has already been deployed to instantiate expertise-driven systems in nutrition (DietNerd), news (NewsNerd), cloud technology (CloudNerd), and space science (SpaceNerd). CustomNerd incorporates specialized knowledge from domain-specific data sources that a user provides and processes the results of searches on those sources. The system consists of five main components: (i) Query Conversion, which converts user input into structured queries for one or more application-specific databases of validated articles; (ii) Extraction, which uses the structured queries to extract relevant data from those external databases; (iii) Filtering and Classification, which identifies and prioritizes relevant documents among those retrieved; (iv) Reliability Assessment, which evaluates the quality of the documents (e.g., statistical validity, absence of conflict of interest); and (v) Response Generation, which produces text outputs based on data sources considered to be relevant and reliable. In addition, the system can (i) generate questionnaires based on a user query, (ii) allow users to upload their own documents, and (iii) work with a variety of Large Language Model (LLM) engines. This paper describes how the system can be used to build an expertise-driven application through a running example, explains the system’s architecture, discusses related work, and evaluates CloudNerd on a 120-question Stack Overflow benchmark against OpenDeepResearch and AutoGPT using semantic textual similarity and RAGAS metrics with paired permutation tests. The results show that the RAG frameworks of CustomNerd and OpenDeepResearch both achieve better semantic fidelity than the agentic AutoGPT. Further, CustomNerd and OpenDeepResearch show complementary semantic advantages. Further, simply concatenating the two answers yields semantic benefits compared to the outputs of either one alone.
Large language model (LLM) applications are rapidly moving from general capability demonstrations to domain-specific engineering applications. Unlike public benchmark evaluations, project acceptance testing emphasizes whether a delivered system can provide correct retrieval, faithful generation, traceable evidence, and...
Feng-Zhi Wei, Tao Xie, Yao Wang et al.· 2026 12th International Conf...· 0 citations
Industrial-Instruction addresses the gap in public instruction-tuning or benchmark datasets built from real industrial technical reports, and releases two parallel versions built by the same pipeline, enabling a direct comparison of open- versus frontier-model data generation.
Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour et al.· 0 citations
OmniExtract is an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks and it can support a comprehensive data extraction including text and tables.
A novel chatbot architecture leveraging OpenAI's GPT-4 model for automated extraction and analysis of repository data is introduced, suggesting that this architecture can make repository data more accessible to technical and non-technical audiences through the production of actionable insights.
Muhammad Jawad Chowdhury, Md. Sakib Khan· 1 citation
By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
A literature-based architectural framework for reliable knowledge retrieval systems that separates external knowledge management from LLM-based reasoning and generation is developed and indicates that reliable LLM deployment should be treated as an end-to-end architectural problem rather than solely a model-performance...
Bharat Kumar Reddy Karumuri· International Journal of Eng...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.