Speaker
Description
FAIR research data originate where research is performed, in laboratories, computational workflows, Electronic Lab Notebooks (ELNs), and collaborative research environments. These environments are designed to support scientific work rather than the publication of reusable research assets. Existing FAIR ecosystems primarily focus on trusted repositories as the entry point for FAIR data, leaving a critical architectural gap between research environments and long term publication infrastructures.
We argue that this missing infrastructure layer is essential because FAIRification is a continuous process rather than a single publication step. Metadata, semantic annotations, provenance information, and relationships between research artefacts evolve throughout a project's lifetime as scientific understanding improves, collaborators contribute additional knowledge, and community standards mature. Freezing metadata together with research data at publication time therefore conflicts with the evolutionary nature of scientific research.
We propose DataHUB, a controlled FAIRification infrastructure that fills this gap combining a Git like collaborative workspace coupled with a publication registry. Within the workspace, researchers collaboratively manage their data together with metadata under version control, enabling continuous FAIRification throughout the research lifecycle. Data, metadata, provenance, and semantic annotations evolve together while remaining fully traceable and reproducible.
The central concept of DataHUB is the separation between the research object and its FAIR representation. Instead of publishing datasets directly, DataHUB publishes a FAIR Research Manifest, an ARC RO Crate inspired, machine actionable representation of the research object. The FAIR Research Manifest contains persistent identifiers, rich metadata, provenance, semantic descriptions, and resolvable references to the associated research artefacts. The published manifest therefore becomes the stable, citable representation of the research object, while the underlying data remain independent of a particular storage location.
This architecture establishes a simple but powerful model consisting of Research Object, FAIR Research Manifest, Repository. Because the FAIR Research Manifest is independent of the physical storage location, datasets can subsequently be transferred to discipline specific repositories, institutional repositories, or other storage infrastructures without modifying the published manifest. Equally important, the same approach supports research communities for which no trusted repository yet exists, for example because a method is newly emerging or because a dedicated repository has not yet been established.
By separating collaborative FAIRification from long term archival storage, DataHUB establishes FAIRification as a controlled, versioned, and collaborative process rather than a one time publication event. The proposed architecture complements existing repository ecosystems, enables continuous metadata evolution, and provides the missing infrastructure layer between research environments and trusted repositories while remaining flexible enough to support future repositories and emerging scientific communities.
The proposed infrastructure is based on a service-oriented model on the PaaS layer of the NFDI Overall Architecture where researchers interact with the consortium’s service layer rather than directly with the underlying technical providers (e.g., storage). This establishes a clear B2B relationship between the consortium and the technical providers. Evolving from our initial CoRDI recommendations to use a modified GitLab and the subsequent opening of the ARChub DataPLANT workspace, the resulting DataHUB integrates seamlessly into the NFDI overall architecture. As a rather basic service it bridges the gap between generic hardware resources and community-specific services.
| background | DataPLANT consortium, common infrastructure, accounting, storage |
|---|