Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.
OBJECTIVE To evaluate the performance and consistency of Large Language Models (LLMs) in core systematic review (SR) tasks and to introduce open-source tools for automated batch processing that provide decision rationales. METHODS We assessed GPT-4o, Kimi-K2, DeepSeek-V3, and DeepSeek-R1 on five SR tasks: title/abstr...