MPG NFDI Meeting

→ Europe/Berlin
MPI for Sustainable Materials

MPI for Sustainable Materials

Düsseldorf
Erwin Laure (MPCDF), Peter Benner (Max Planck Institute for Dynamics of Complex Technical Systems)
Description

The next MPG NFDI workshop will take place from 27.10 12:00 - 28.10. about 16:00 in Düsseldorf, hosted by the MPI for Sustainable Materials. The key theme will be to discuss how the various NFDI Consortia, MPG researchers are involved in, deal with AI, particularly how data is being prepared for training purposes and other uses of AI in the projects. Of course, other topics are possible, too - please let us know if you have any suggestions.

We are soliciting in particular contributions discussing how Consortia are employing AI technologies and make their data available for AI training, however other topics are welcome as well. 

 

Registration
Registration
  • Tuesday 27 October
    • 12:00 → 13:00
      Light Lunch 1h
    • 13:00 → 13:15
      Welcome 15m
    • 13:15 → 14:00
      AI usage in the second phase of NFDI4Earth 45m

      MPI-BGC is a co-applicant in the second funding phase of NFDI4Earth, which starts in October 2026. We present a brief overview of the project and our role in it. We then focus on which services are built around AI. Specifically, we show how Foundation Models are developed in geosciences and how NFDI4Earth can improve their reusability.

    • 14:00 → 14:45
      Knowledge Graph construction for Bioimaging 45m

      Semantic web resources or Knowledge Graphs have emerged as a preferential source of training data for large language models and generative artificial intelligence. Domain specific Knowledge Graphs are currently being constructed by various NFDI consortia to enhance AI readiness and interoperability of data and metadata within and across these domains.

      Here, we outline the construction of Knowledge Graphs from a bioimage data repository. We employ an ontology-based mapping mechanism between the underlying relational database and the semantic web layer, offering a SPARQL endpoint for semantic queries.

      Recently, we have augmented our solution with GeoSPARQL capabilities, allowing query clients to find, access, and relate images based on geolocation coordinates. This provides a concrete example of how domain-specific metadata can be made discoverable and accessible to users and AI agents beyond the bioimage community.

      Furthermore, we are exploring the federation of Knowledge Graphs across
      independent bioimage repositories. Such a federated semantic layer could
      enable AI agents to discover and access distributed bioimaging data through a common query interface, without requiring knowledge of the underlying repository implementations.

    • 14:45 → 15:15
      Coffee 30m
    • 15:15 → 16:00
      AI-ready experimental data in materials science with FAIRmat 45m

      AI applications in materials science and catalysis increasingly depend on access to large, well-structured and semantically meaningful experimental datasets. In this contribution, we will discuss how FAIRmat addresses these challenges through community-agreed data standards, semantic data models, and the NOMAD ecosystem.1,2 Examples from activities at the MPI for Chemical Energy Conversion (MPI CEC) will illustrate these approaches. A particular focus will be on transforming heterogeneous experimental data into interoperable datasets that can be used for AI training and other machine-learning applications. We will present examples from experimental materials-science domains, including spectroscopy and microscopy, and discuss how NOMAD supports the generation, contextualization, and reuse of AI-ready data. We will also outline planned developments for FAIRmat’s second funding period, including the integration of AI workflows with the FAIRmat/NOMAD ecosystem and planned AI applications at MPI CEC that build on the resulting data models and AI infrastructure.

      (1) Scheidgen et al., (2023). NOMAD: A distributed web-based platform for managing materials science research data. Journal of Open Source Software, 8(90), 5388, https://doi.org/10.21105/joss.05388
      (2) Schumann, J.; Näsström, H.; Götte, M.; Himanen, L.; Moshantaf, A.; Scheidgen, M.; Márquez, J. A.; Draxl, C.; Trunschke, A. Enabling open and FAIR catalysis data with standardized data structures. Nature Catalysis 2026, 9 (3), 225-229. https://doi.org/10.1038/s41929-026-01508-9.

    • 16:00 → 16:45
      PUNCH and Beyond (over video) 45m

      tbd

    • 16:45 → 17:30
      MatWerk 45m

      tbd

    • 18:30 → 21:30
      Dinner 3h
  • Wednesday 28 October
    • 09:00 → 09:45
      AI in the Photographic Collection for data curation and cataloguing (over video) 45m

      The Fotothek of the Bibliotheca Hertziana has been experimenting with AI tools to produce datasets, to improve applications and to support scientific cataloguing with different results on different tasks. I would like to report on this and detail successes and failures, including some more detailed testing done for inscriptions part of the Art collection of the institute.

    • 09:45 → 10:30
      LLM @ Text+ (preliminary) 45m

      TBD

    • 10:30 → 11:00
      Coffee 30m
    • 11:00 → 11:45
      Automated workflows for scientific computing and AI training through MaRDI services 45m

      Modern scientific workflows increasingly depend on research software, yet numerical experiments often face challenges in reliability, robustness, and interoperability. Although novel algorithms continue to emerge, their adoption across disciplines is frequently limited by the difficulty of integrating them into established workflows.
      Research data are often insufficiently annotated, inconsistently structured, or difficult to access, and hence not directly usable for AI training purposes. Also, the lack of standardized benchmarks hinders transparent and reproducible performance evaluation. In multi-step research processes, data, tools, and domain-specific knowledge are also rarely captured in forms that can be translated into automated workflows. To address these challenges, MaRDI provides a standardized specification and services for data and metadata management, benchmarking, and workflow representation across mathematical domains. Designed for extensibility, customizability, and reproducibility, MaRDI services support a flexible, domain-independent approach to scientific workflows within and across mathematical communities.

    • 11:45 → 12:30
      Base4NFDI 45m

      Base4NFDI is a joint initiative of all discipline-specific 26 Consortia within the National Research Data Infrastructure (NFDI) to foster and establish reliable NFDI-wide basic services for FAIR research data management. Here, I will provide an update on current activities (development projects supported, incubators, user conference, ...) as well as ideas and discussions about a potential future as NFDI moves into its second phase.

      https://base4nfdi.de/

    • 12:30 → 13:30
      Lunch 1h
    • 13:30 → 14:15
      Accounting und IAM4NFDI als Grundlage für Ressourcenmonitoring in der NFDI 45m

      Neben der Frage, wie Konsortien Daten für KI-Training aufbereiten, stellt sich die Frage nach der zugrundeliegenden Infrastruktur: Wie wird Ressourcennutzung über Konsortiengrenzen hinweg erfasst, zugeordnet und abgerechnet?

      Dieser Beitrag stellt den aktuellen Stand von Accounting4NFDI vor: die Erfassung und Aggregation von Ressourcennutzungsdaten aus verschiedenen Konsortien und Infrastrukturen, mit dem Ziel einer konsortienübergreifend einheitlichen und nachvollziehbaren Abrechnungsgrundlage. Wir gehen auf die technische Umsetzung des Ressourcenmonitorings ein sowie auf offene Fragen bei der Harmonisierung heterogener Ressourcentypen (Compute, Storage, ggf. spezialisierte KI-Ressourcen) über verschiedene Betreiber hinweg.

      Eine Voraussetzung für belastbares Accounting ist eine verlässliche Identitätsgrundlage: IAM4NFDI liefert die föderierte Zuordnung von Nutzenden zu Organisationen, Projekten und Rollen, auf der Ressourcennutzung erst zugeordnet werden kann. Wir skizzieren den aktuellen Stand von IAM4NFDI und die Schnittstellen zwischen IAM und Accounting.

      Der Beitrag ordnet diese Infrastrukturarbeit in den Kontext der aktuellen KI-Diskussion ein, ohne KI selbst in den Mittelpunkt zu stellen: Ressourcenintensive KI-Anwendungen setzen funktionierendes Accounting und IAM voraus, unabhängig davon, wie die KI-Nutzung im Detail aussieht.

    • 14:15 → 15:45
      The Future of NFDI and The Role of MPG 1h 30m
    • 15:45 → 15:55
      Closure & Farewell 10m