Speakers
Description
A decade has passed since the publication of the FAIR Principles, and evaluating how data-sharing practices have evolved in scholarly publishing is informative for shaping sustainable Research Data Management (RDM) infrastructures. Within the German National Research Data Infrastructure (NFDI) and its Basic Services, understanding where policy mandates succeed or fall short is essential to avoid redundant developments and build targeted, machine-actionable services. However, large-scale empirical evidence on whether research outputs are really becoming FAIRer or merely superficially complying with relevant mandates remains sparse.
To address this gap, we developed an automated, reproducible eight-stage ingestion and analytical pipeline in Python to conduct a large-scale longitudinal corpus study of over 5.7 million biomedical and life-science articles from the PubMed Central Open Access Subset (PMC-OA), published between 2011 and 2025. The pipeline extracts and normalizes Data Availability Statements (DAS), ethics disclosures, repository URLs, accession numbers, and supplementary file links from JATS XML structures. Using rule-based text classification, persistent identifier parsing, automated F-UJI FAIR assessments, and JATS for Reuse markup compliance checks, we evaluated articles against a standardized proxy FAIR score. Temporal trends before and after 2016 were analyzed using segmented linear regression and interrupted time-series modeling, with stratification by discipline, publishing model, collaboration type, and author demographics.
Our analysis reveals a complex shift in data-sharing behaviors. An adoption versus quality paradox emerges: while explicit DAS adoption surged dramatically from under 3% in 2012 to over 65% in 2024, this growth was matched by a parallel surge in the "available on request" anti-pattern, which grew from near 0% to over 18%. This indicates that mandates alone do not guarantee machine-actionable reusability. Furthermore, substantial disciplinary disparities persist. Structural biology, genomics, and bioinformatics consistently achieved the highest FAIR scores, driven by mature domain-specific repositories like NCBI GEO, GenBank, and PDB, whereas clinical research and public health usually scored considerably lower. Regarding publishing models, articles in Gold Open Access journals generally outperformed subscription titles, and while high-tier repository links grew steadily, generic or non-standardized links still constitute a substantial proportion of shared URLs. Finally, international collaborations scored higher than domestic ones, and a persistent gender gap emerged, with female first-authored papers showing significantly less post-2016 FAIR score improvement.
These empirical trends provide actionable insights for Base4NFDI and cross-cutting RDM needs. The prevalence of non-machine-actionable statements underscores the necessity of linking structured DAS generators directly into DMP4NFDI workflows and Jupyter environments. The persistent use of generic URLs highlights the urgency of PID and terminology harmonization via PID4NFDI and TS4NFDI. Moreover, observed domain gaps show where targeted onboarding is required, while automated compliance layers like F-UJI and JATS4R demonstrate how continuous FAIR monitoring can be integrated into national basic services.
| background | Jupyter4NFDI, MaRDI, KGI4NFDI, SeDOA |
|---|