24–25 Nov 2026
Europe/Berlin timezone

From FAIR Publication to FAIR Exploration: A Repository-to-Knowledge Graph Integration Architecture

24 Nov 2026, 16:54
1m
Otto-Braun-Saal (Stabi Berlin, Potsdamer Straße)

Otto-Braun-Saal

Stabi Berlin, Potsdamer Straße

Posters Posters

Speakers

Yuliia Dikova (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart) Volodymyr Kushnarenko (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart)

Description

From FAIR Publication to FAIR Exploration: A Repository-to-Knowledge Graph Integration Architecture
Authors:
Yuliia Dikova1, Volodymyr Kushnarenko1, Preston Rodrigues1, Nadiia Huskova1 and Thomas Boenisch1
yuliia.dikova@hlrs.de
1High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart

Abstract:
FAIR principles have improved the publication and long-term accessibility of research data. Repository platforms such as Dataverse enable researchers to publish datasets, assign persistent identifiers, preserve multiple versions, and make research outputs findable and accessible. In several scientific domains, repositories increasingly host not only descriptive metadata but also domain-specific semantic metadata, typically represented as RDF.
Publishing semantic metadata, however, does not automatically make these metadata easy to explore across multiple publications. While repositories allow researchers to discover individual datasets, answering scientific questions such as "Have similar experiments already been published?" or "Which publications contain comparable experimental conditions?" still requires downloading RDF files from multiple datasets, loading them into a local triple store, and manually combining them before any meaningful analysis can begin. As the number of published datasets grows, this process becomes increasingly repetitive and difficult to maintain.
In this contribution, we present an integration architecture that extends FAIR publication towards FAIR exploration without changing existing publication workflows. Rather than introducing another publication repository, the proposed approach connects an existing repository with a domain-specific semantic service that continuously synchronizes published semantic metadata into a shared knowledge graph. In our implementation within NFDI4Cat, this architecture is realized by integrating Repo4Cat repository with semantic platform SemMan4Cat .
SemMan4Cat periodically detects newly published or updated datasets through the Dataverse API, retrieves the published semantic metadata together with the associated publication metadata, validates the incoming RDF, manages versions, and incorporates the semantic content into a shared triple store. The repository remains the authoritative source for publication, preservation, citation, and long-term access, while the semantic platform complements these capabilities by providing services that general-purpose repositories typically do not offer.
Once synchronized, semantic metadata from multiple publications become immediately available for cross-publication querying, visualization, and semantic exploration. Researchers no longer need to repeatedly download and locally integrate RDF files in order to compare published experiments. Instead, they can explore the complete collection of synchronized semantic metadata and, whenever relevant results are identified, directly navigate back to the corresponding dataset, metadata files, and persistent identifiers stored in the repository. Figure 1 summarizes the proposed repository-to-knowledge graph integration architecture, highlighting the continuous synchronization of semantic metadata from Repo4Cat into SemMan4Cat and the navigation back to the original repository publications.

Figure 1 - FAIR Research Lifecycle

The integration is implemented using the Dataverse External Tools framework. Researchers can launch SemMan4Cat directly from a published Repo4Cat dataset without requiring a separate account for publicly available data. The selected publication serves as the starting point for semantic exploration across all synchronized publications, creating a seamless transition from repository search to knowledge graph exploration.
Our contribution presents the architecture and the first implementation of this repository-to-knowledge-graph integration. Rather than replacing existing repository infrastructures, the proposed approach demonstrates how reusable semantic services can extend them with advanced discovery and exploration capabilities while preserving repositories as the single authoritative publication platform. The proposed architecture is transferable to research communities that publish semantic metadata and require scalable cross-publication discovery while preserving repository-based publication workflows.

Keywords: Repository-to-Knowledge Graph Integration, Research Data Repositories, Semantic Metadata, Cross-Publication Discovery, FAIR Data

background NFDI4Cat

Author

Yuliia Dikova (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart)

Co-authors

Volodymyr Kushnarenko (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart) Preston Rodrigues (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart) Nadiia Huskova (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart) Thomas Bönisch (High-Performance Computing Center Stuttgart (HLRS), University of Stuttgart)

Presentation materials