Standards of standards: Relevance of harmonising workflows for data integration and analysis
Abstract
The continuous increase in the amount of available data over recent decades has made it very clear that programming and bioinformatics tools are essential to handle the many datasets and records, which are available in ecology. Without such tools including open scripting languages like R or Python, data preparation and statistical analyses would be impossible in the extent as applied today. The use of scripts for conducting data analyses is also a big step towards reproducible, transparent, and open science as even tiniest steps of data processing can now be easily documented, which would not be described in standard research publications with limited space for methods sections. As another consequence of the increasing data volume, the consideration of a large amount of data has become standard in scientific publication and is also expected by the peer review community. Thus, in many cases workflows for data harmonisation and integration need to be developed before the actual analysis. Currently, such workflows are developed mostly in isolation within individual research projects by individual researchers. Workflow development involves many, often small decisions of how to handle certain records or how to conduct a certain step of the whole data processing workflow. Although often somewhat documented by providing code these many decisions are usually made based on the knowledge, capacity, and interest of the individual researchers without following clear standards. Standards have been developed for various aspects of data processing. Prominent examples of standards in biodiversity research are the development of accessible taxonomies or controlled vocabulary such as Darwin Core. However, the application of these standards is often not straightforward and not done consistently across approaches. In addition, many other aspects of such workflows do not have any standard or existing standards are not applied. This involves comparatively easy things like clear and unique geographic names but also the handling of uncertainty in the provided information or decisions about non-matching taxonomic names. By highlighting currently common approaches, their challenges and limitations, I will make the point that we therefore need standards for conducting the harmonisation and integration of data and for developing respective workflows. Guidance and recommendations of how workflows for data harmonisation should be developed are required also to improve interoperability of datasets. The development of standards for standardisation following the FAIR data principles would also provide an opportunity to more easily compare and share code. Making such code as generic and flexible as possible would also provide an opportunity to set new standards as other researchers, who are less familiar and interested in that topic, may use the code or approaches for their projects, which finally will make scientific products more comparable and code reusable in an easier way.