Accelerating EICAT Assessments Using AI & Evaluating inter-annotator/LLM agreement
Abstract
AI and Large Language Models (LLMs) are increasingly applied to modern information processing, and ecology and biodiversity are no exception. The OneStop (2025) project seeks to minimise the introduction, establishment and spread of terrestrial invasive alien species. As part of this project, we are investigating how LLMs can accelerate the Environmental Impact Classification for Alien Taxa (2026) assessment process (Hawkins 2015). Prior work has explored smaller, locally run encoder/decoder models for EICAT impact classification (Brinner and Zarrieß 2025). We build on this by investigating whether larger, general-purpose LLMs offer greater flexibility across the full EICAT schema. Our system assists human annotators by ingesting PDF documents, converting them to a structured format, and applying LLMs to identify invasive species impacts, whilst keeping humans-in-the-loop to ensure consistency and reliability. The system extracts evidence text, EICAT classification level, impact mechanism and identifies the impacted native species. To benchmark the system, we are running an inter-annotator evaluation with 22 annotators, all of whom are scientists in the domain and around half of whom have experience applying the EICAT assessment protocol. Papers for evaluation are drawn from existing EICAT assessments in the Global Invasive Species Database (GISD). The evaluation uses a Balanced Incomplete-Block Design with a core set of papers reviewed by all annotators. Annotators are asked only to read the EICAT guidelines before starting — intentionally minimal, to help surface ambiguities in the guidelines themselves. We use Cohen's kappa (Cohen 1960) to measure pairwise agreement on impact identification and Krippendorff's alpha (Krippendorff 2011) across the full multi-label classification scheme. We also perform automated evaluation on a wider set of GISD texts, comparing single-annotator extractions against system outputs using precision, recall and F1 scoring across impact identification, mechanism and severity classification. We trial multiple LLMs via AWS Bedrock (GPT-OSS, Amazon Nova, Claude, Llama), assessing whether lower-cost models can perform as reliably as human assessors. The tool (Suppl. material 1) is brought together in a web-based portal covering the full EICAT assessment workflow. We also plan to evaluate ASReview (Schoot 2021) and LatteReview (Rouzrokh and Shariatnia 2025) to understand if they can speed up abstract screening at the literature review stage. The portal follows FAIR data principles (Wilkinson 2016) and is designed to remain available beyond the life of the OneSTOP project, either via further funding sources or as an open-source, self-hosted solution. A preliminary report will be available in July 2026, with a full cost-benefit analysis by the end of the year, providing practical guidance for organisations considering whether AI-assisted EICAT assessment is worth the investment.