close icon

Advancing research data tracking with the Open Science Monitoring Initiative

September 23, 2025   |   By: Make Data Count

Laetitia Bracco

OSMI

DOI: 10.60804/8v2k-cm37

 

The Open Science Monitoring Initiative promotes the open science monitoring principles and adoption of practices to understand the progress of open science. Make Data Count is an active participant in OSMI, through its Working Group focused on open science monitoring with scholarly content providers. We had a conversation with Laetitia Bracco, member of the OSMI Coordination Committee, about OSMI’s activities and collaboration opportunities between OSMI and Make Data Count to advance data monitoring and evaluation practices.

Can you tell us a bit about your role at the Université de Lorraine, and your involvement with open data and open software as part of the French Open Science Monitor project.

Currently Deputy Head of the Libraries Research Support services, I have been working at the Université de Lorraine since 2019. I am especially in charge of research data and bibliometrics. Université de Lorraine was the first university to reuse the code and indicators of the French Open Science Monitor (FOSM) at a local level, in 2020. That’s why in 2021, the French Ministry of Higher Education and Research asked us to manage the project to extend the FOSM to research data and software. I have been project manager for this area of the FOSM for four years now.

The goals of the project are to track down data and software mentions within scientific publications, characterising those mentions to analyse open science practices, and to evaluate the use of data repositories for sharing and opening research data in France.

You are also part of the Coordination committee for the Open Science Monitoring Initiative (OSMI). Can you tell us a bit about OSMI, its focus, and current work on open science?

The Open Science Monitoring Initiative was launched officially in December 2024. Its main purpose is to encourage the adoption of open science monitoring principles and to promote their practical implementation. Indeed, as stated in May 2023 by the G7 Science and Technology Ministers, an international collaboration is needed to “inspire a framework for open science monitoring”. This framework had to be defined and this is the whole purpose of the Principles of Open Science Monitoring, published in July 2025 after an international consultation led by OSMI and facilitated by UNESCO.

The founding members of OSMI are the French Ministry of Higher Education and Research, Université de Lorraine, UNESCO, Inria, SPARC Europe, PLOS and the Berlin Institute of Health (BIH) at Charité. But as of today, OSMI counts nearly 200 experts from all areas of the world among its members.

OSMI also comprises four working groups currently engaged in scoping the needs of open science monitoring, understanding the worldwide landscape of open science monitoring, helping scholarly content providers monitor open science, and identifying the tools to analyse open scholarly outputs.

Open data is a key aspect of open science. Make Data Count’s goals to advance a better understanding of data usage and OSMI’s activities on practices and trends to monitor open outputs converge for open data. How do you see Make Data Count’s work aligning with and supporting OSMI’s mission and activities?

Indeed, FAIR and open data is a key aspect of open science monitoring. As stated in the Principles, it is paramount to “emphasize quality and integrity, equity, inclusion, collective benefit, fairness, inclusivity, and recognition of diverse open science practices, outputs, and outcomes”, which implies among other things that monitoring publications only is not enough to fulfil monitoring of open science. 

A better understanding of the ways in which research data is shared, published, cited and reused is one of the key aspects – and not an easy one, as research data are a much more nuanced object than publications. It also takes a village to properly share and open data: researchers are at the forefront, but publishers, librarians, developers etc also play a key role. For this, we are very grateful to Iratxe Puebla, director of Make Data Count, to be one of the Co-chairs of OSMI working group 3, “Open science monitoring with scholarly content providers”.

Make Data Count is developing the Data Citation Corpus, a large aggregate of data citations, and seeking to leverage developments in large language models to identify and expose data-article connections. What opportunities do you envision for these resources to support future efforts in open data monitoring? 

I think that large language models (LLMs) are the future of text and data mining, especially for research data. Indeed, data is still rarely included as a structured citation in the reference list, and is not always published in repositories with complete associated metadata. Those elements make data use or reuse very difficult to capture. Finding mentions of data production and use in the full text of scientific publications is therefore another way to pursue the identification of datasets. 

Today, some open source tools already exist, but all have some limitations and there are also disciplinary discrepancies. I do think that LLMs might be an answer to better identify explicit (like “we use the dataset with ID X stored in the Y data repository”) as well as implicit (such as “to obtain these results, temperature records were analysed”) data mentions. One of the biggest challenges will be to create trust around those tools, while implementing them in a way that is sustainable and has the lowest environmental impact possible.

In your view, what are the most pressing needs in relation to data monitoring and data evaluation? What are technical, social or cultural advances you would like to see?

Skills, tools, infrastructures: I think there are multiple needs with regard to data monitoring. As more and more libraries are engaged in the process of open science monitoring, we need simple solutions that they can deploy to monitor research data. 

Firstly, to track the datasets generated by an institution. In France, the French Ministry of Higher Education and Research developed, within the FOSM project, a tool dedicated to finding datasets in repositories: https://works-magnet.esr.gouv.fr/datasets/search. Secondly, to analyse the full text of publications to identify associated datasets. Last but not least, it is essential to develop a culture of data citation among researchers. For transparency and appropriate attribution, we need to emphasise that best practices require that the data used to obtain a scientific result is cited, with equivalent citation standards as those for journal publications.

Data evaluation is also a key aspect. If we want researchers to publish their data, this practice has to be rewarded. This is why I have great expectations for CoARA, and other initiatives like Make Data Count’s “Implementing data evaluation at institutions” working group, to advance inclusion of data in research evaluation.

We cannot make progress without giving the researchers some incentive – other than working for the greater good – to share their data.