Skip to content
Preprint

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

Aug 2026 · 0 citations · 52 references
Computer Science

TL;DR

Results indicate that SciDSK improves how agents locate and understand scientific datasets, providing a stronger foundation for actionable scientific data use.

Abstract

Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for reliable dataset discovery and interpretation, constraining their effective use in scientific workflows. This limitation arises because agents must search across heterogeneous repositories and reconstruct dataset-specific semantics and operating procedures from documentation designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, task-specific usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. On the query retrieval benchmark, Agent-SciDSK achieves 80.77% Hit@1, exceeding Agent-Raw by 9.62 percentage points. Across controlled interpretation cases, the SciDSK condition satisfies 23 of 24 assessment criteria, compared with 22 under the web-page condition. These results indicate that SciDSK improves how agents locate and understand scientific datasets, providing a stronger foundation for actionable scientific data use.

View source

Similar papers

Book Open access Jul 2026

Toward a Trustworthy and Accessible Scientific Data Workflow Platform with StreamCI

Scientific research workflows increasingly involve not only structured streaming data but also raw artifacts and derived products, requiring platforms that provide trustworthy data protection and accessible interfaces beyond simple ingestion and storage. We present extensions to StreamCI, a cloud-based streaming data m...

Jaewoo Shin, Mehak Jain, S. Jha et al. · 0 citations
Jul 2026

SciDataSailor: Deep Scientific Data Exploring

This work presents SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation and presents SciDataSailor, a framework for synthesizing tool-interactive trajectories as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms.

J. Rao, Yicheng Qiu, Chi Zhang et al. · 1 citation
Preprint Aug 2026

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explore. Although many datasets are publicly available, using them often requires familiarity with repository organization, data formats, multiresolution structures, and visual...

Aashish Panta, Hugo Lee, G. Scorzelli et al. · 0 citations
#large language models Open access Sep 2026

From Data Quality to Quality of Agentic Data Use: A Conceptual Framework for Agentic Data Engineering

Large language models and AI agents are extending data-engineering automation beyond isolated artifact generation toward end-to-end processes in which agents interpret requirements, select data, generate transformations, invoke tools, validate results, and communicate analytical outputs. This shift introduces risks tha...

Ania Cravero, Jorge Díaz, Zihao Xiao · 0 citations
Review Open access Jul 2026

FRED enables standardized FAIR metadata generation and management for omics research

Scientific research relies on transparent dissemination of data and its associated interpretations, including raw data, metadata, experimental design, and data processing details. Production and handling of research data represents an ongoing challenge, extending beyond publication into individual facilities, institute...

Jasmin Walter, C. Kuenne, Noah Knoppik et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.