24–25 Nov 2026
Europe/Berlin timezone

ReproScore-Any: Turning a Reproducibility Study into a Basic Service

24 Nov 2026, 16:31
1m
Otto-Braun-Saal (Stabi Berlin, Potsdamer Straße)

Otto-Braun-Saal

Stabi Berlin, Potsdamer Straße

Posters Posters

Speaker

Vasundhara Shaw (FIZ-Karlsruhe)

Description

Computational notebooks are routinely deposited alongside publications, but deposition does not necessarily imply executability. In a previous study of biomedical publications indexed in PubMed Central, 15,817 Python notebooks with declared dependencies were rerun automatically: 7.6% completed without error, and only 5.6% reproduced the results originally reported [1]. That result raised two questions we will address in this talk: (1) does it hold outside biomedicine, and (2) can computational reproducibility assessment be offered as a service rather than repeated as a study in every field?
To answer the first, we ported the pipeline to astrophysics, replacing PubMed Central with the NASA ADS literature index. Only one component had to change: the discovery layer, which links publications to their code repositories. Everything after that, cloning the repository, rebuilding the environment, running the notebooks, and classifying the errors, worked unchanged. The field-specific part of computational reproducibility assessment is smaller than we had expected.
We answered the second by building ReproScore-Any: give it a repository URL, and it returns a Reproducibility Score with sub-scores for execution success, data availability, documentation, code quality, and structure [2]. Discovery is now pluggable, so a consortium can supply its own literature source and reuse the rest. The service runs on the NFDI JupyterHub, with two additional developments underway: deployment as a containerised service on Jupyter4NFDI infrastructure and integration into the KGI4NFDI infrastructure.
We will present results from astrophysics and computational biology repositories, including the score distributions and the classes of failure that recur across both fields. We close with the question we would most like to discuss with this audience: whether such a service belongs in the NFDI as an author-facing pre-submission check, a repository-side quality signal, or an instrument for measuring computational reproducibility at scale.
Code, dashboard, and demo are openly available [3].
[1] S. Samuel, D. Mietchen, Computational reproducibility of Jupyter notebooks from biomedical publications (2024) GigaScience 13:giad113, https://doi.org/10.1093/gigascience/giad113
[2] Samuel, S., Mietchen, D., Kim, J., Ahmed, W. and Gaedke, M., 2026. ReproScore: Separating Readiness from Outcome in Research Software Reproducibility Assessment., TPDL 2026. arXiv preprint arXiv:2605.13275.
[3] Link to repository for testing for reproducibility https://vasundharashaw.github.io/ReproScore-Any/

background Jupyter4NFDI

Author

Vasundhara Shaw (FIZ-Karlsruhe)

Co-authors

Dr Moritz Schubotz (FIZ-Karlsruhe) Dr Daniel Mietchen (FIZ-Karlsruhe) Sheeba Samuel (GNOI)

Presentation materials