Speaker
Description
A key challenge in achieving the interoperability facet of the FAIR data principles is the consistent identification of researchers, chemicals, paintings, and other entities relevant for NFDI. Often, this means choosing the correct standard uniform resource identifier (URI) or compact URI (CURIE) for an entity from an ontology, controlled vocabulary, PID service, or database.
This presents a challenge when multiple CURIEs or URIs can be constructed for the same entity, which is compounded by the independent evolution of tools and services relevant for different NFDI consortia. For example, the entry for water in the Chemical Entities of Biomedical Interest (ChEBI) can be identified by either the URIs https://www.ebi.ac.uk/chebi/CHEBI:15377 or http://purl.obolibrary.org/obo/CHEBI_15377 or by the CURIEs chebi:15377, CHEBI:15377, or CHEBIID:15377. An organization-wide policy is required to determine and communicate which is correct.
The NFDI has not yet adopted an actionable, organization-wide policy for standardizing CURIEs and URIs. We present the Semantic Farm (previously called The Bioregistry) on behalf of the Section Metadata Working Group for Ontology Harmonization and Mappings as a pre-existing, mature solution for the standardization of CURIEs and URIs. It acts as a centralized index of metadata for resources that mint identifiers that can be readily adopted on the NFDI-level and beyond.
Importantly, the Semantic Farm is a foundational service that can directly support complementary base services (PID4NFDI, TS4NFDI, and KGI4NFDI), sections, their respective working groups, and researchers in NFDI consortia towards improving interoperability. We highlight several existing applications of the Semantic Farm within NFDI:
- TS4NFDI uses both the flagship instance of the Ontology Lookup Service (hosted by the European Bioinformatics Institute) and the TIB Terminology Service instance, which both use the Semantic Farm for URI compression, CURIE expansion, and generation of web links for database cross-references.
- Section Metadata Working Group for Ontology Harmonization and Mapping uses the Semantic Farm capture ontology lists used by each consortium (https://semantic.farm/nfdi).
- The Semantic Farm supports the construction of knowledge graphs with standardized URIs such as in Section EduTrain's DALIA platform for open educational resources and Section International Engagement Working Group for Landscaping and Outreach's bibliometric knowledge graph. Standardization facilitates the integration of external data from ORCiD, ROR, CORDIS, and Wikidata.
- The Semantic Farm's codebase is used by the LinkML runtime, which has growing adoption across consortia such as NFDI4Chem, NFDI4Cat, and GHGA. It can be further used to standardize the prefix maps in LinkML schemas.
- The Semantic Farm supports consortia like NFDI4Chem that must standardize and ultimately harmonize a variety of ontology, database, and structural identifiers like InChI.
Further information on the Semantic Farm's KPIs can be found here. We aspire to submit Semantic Farm to Base4NFDI Service Support Track and to support its integration in more NFDI services and usage in data resources.
| background | Section Metadata WG on Ontology Harmonization and Mapping |
|---|