Skip to content

Author

Abdullah Keleş

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

How Often Do Large Language Models Agree with Each Other—And with the Truth? A Consensus- and Complexity-Stratified Analysis of Data Extraction for Neuroimaging AI

Background: The reliable integration of large language models (LLMs) into neuroimaging data extraction workflows remains unresolved. Prior benchmarking shows that exact-match accuracy underestimates LLM extraction performance, but whether inter-model consensus and variable complexity can guide automation remains unclear. We evaluated whether inter-model consensus can serve as a confidence signal for human–artificial intelligence (AI) extraction and can guide complexity-stratified workflow triage. Methods: Four frontier LLMs were queried via OpenRouter with an identical zero-shot structured prompt to extract 22 predefined variables from 91 peer-reviewed neuroimaging AI articles, yielding 2002 article–variable items per model. Variables were stratified a priori into low- (n = 7), medium- (n = 8), and high-complexity (n = 7). Performance was compared with an expert reference using exact-match and semantic-equivalence accuracy. Item-level consensus and five triage strategies characterized the efficiency–accuracy trade-off. Results: Semantic-equivalence accuracy converged to 80.5–83.4% across models despite approximately ten percentage-point exact-match differences. Unanimous 4/4 consensus occurred in 45.6% (910/1994) of items, with exact-match accuracy of 85.8%, rising to 95.3% after semantic normalization; however, 14.2% still failed to match the reference. Reliability was complexity-dependent: 96.6% for low-complexity variables, 73.2% for medium-complexity variables, and 38.1% for high-complexity variables. A hybrid strategy auto-accepting 4/4 items and routing 3/4 items to rapid verification reduced estimated review effort by approximately 59%. Conclusions: Inter-model consensus is useful, but it is incomplete and depends on variable complexity. We show that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.

Nafiye Şanlıer, Umid Sulaimanov, Ariorad Moniri et al. · 0 citations