It is found that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification.
Abstract
We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.
In recent times there is considerable unease about the way Artificial Intelligence (AI) powered by large language models is used to do astrophysics, and, assuming that these AIs will get far better in future, several alarming situations are also discussed. In a ``white paper,''Hogg argues that extreme cases like allowi...
The first sentence of this abstract--and the introduction to these proceedings--was authored by a human, but the bulk of this document was generated by an agentic AI system. In this talk, I take stock of machine learning (ML) for the LHC physics program over the twelve months from May 2025 to May 2026. The corpus is th...
Generative AI (GenAI), such as ChatGPT, has emerged as a powerful collaborator in numerical problem-solving across scientific domains. However, its reliability in addressing novel physics problems—those beyond the standard undergraduate curriculum—remains largely unvetted. We present a case study investigating the el...
Nam-Jung Kim· Journal of Physics Communica...· 0 citations
This study aims to comparatively examine the potential of four widely used free AI chatbots (ChatGPT, DeepSeek, Gemini, Grok) to spread misconceptions about high school-level physics. Free versions of ChatGPT, DeepSeek, Gemini, and Grok's chatbots were selected as AI tools. As fundamental physics concepts, the concepts...
Hasan Şahin Kızılcık, Nuray Önder Çelikkanlı, A. Tuysuz· International Journal of Edu...· 0 citations
Generative AI (GenAI) promises a disruptive impact across multiple research and operational activities, including the domain of astrophysics and the development of astronomical instrumentation. This rapidly evolving tech nology requires a cross-disciplinary, systemic approach to ensure effective integration into astron...
G. Capasso, C. Arcidiacono, Marco Baldini et al.· Astronomical Telescopes + In...· 0 citations
Modern astrophysical research requires the integration of rapidly expanding scientific literature, heterogeneous observational data, and increasingly complex physical models. We introduce AstroGenesis (https://astrogenai.com), a domain-specific multi-agent AI framework that integrates literature retrieval, multiwavelen...
N. Sahakyan, M. Khachatryan, A. Mahabal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.