Research Gap Discovery Dataset (Portuguese): LLM-Generated Limitations and Research Gaps from 1,724 SciELO Brazil Articles
Abstract
This dataset contains 1,724 Portuguese-language scientific articles from SciELO Brazil (published 2021-2024). For each article, a large language model read the title and Portuguese abstract and wrote, in Portuguese, one sentence each on the article's main limitation (Limitação), an open research gap for future work (Lacuna de pesquisa), and why that gap matters (Importância). The articles span many disciplines, including health sciences, agriculture, education and the social sciences. How the data was gathered: articles were retrieved from SciELO Brazil through the ArticleMeta API. Each title and abstract was sent to mistral-small-latest (Mistral AI, 1,458 articles) or qwen/qwen3.8-27b (Groq, 266 articles) with the same Portuguese prompt (max_tokens 300). The Generator Model column records the model for each row. Abstracts were cleaned (HTML codes, leading "RESUMO" label), and 10 articles with English abstracts were removed. Fields: Serial, Paper ID, Title, Year, Month, Keyword, Abstract, Limitation, Research Gap, Importance, Gap Refers To Own Work (heuristic flag), Generator Model and Split (train/validation/test, 80/10/10, stratified by model, seed 42). How to use it: the data can be used to study research gaps in Portuguese-language science, and to train or evaluate models that write limitations and research gaps from abstracts, including cross-lingual work together with the companion English dataset (doi: 10.17632/px9xd7tw8n.2). Use the Split column for comparable results. Interpretation: the text fields are inferred by an LLM from abstracts and have not been validated by human experts. See README.md for the exact prompt, column definitions and known limitations.