Community perspectives on the Data Citation Corpus
November 26, 2024 | By: Make Data Count
https://doi.org/10.60804/fkfy-1162
Make Data Count is a community hub for the development of metrics that can help us understand how data is used in research and policy activities. Several Make Data Count community members have given us their insights on the Data Citation Corpus, our project to build a central open aggregate of data citations. In this post, they share more about their interest in the Corpus, and why they believe it will be a valuable resource for funders and institutions
Thank you to Francisco Silva-Garcés, Co-fundador Fundación Openlab Ecuador, Coordinador de la Red de Investigación de Conocimiento, Software y Hardware Libre; Ricardo Hartley Belmar, Universidad Central de Chile & Data Observatory Foundation; Chokri Ben Romdhane, CNUDST-Ministry of High Education; and Barbora Bieliková, Slovak University of Technology, for sharing their perspectives.
Why are you interested in the Data Citation Corpus? What uses do you envisage for the Corpus as part of your activities?
Francisco Silva-Garcés:
I see the Data Citation Corpus as a source of open information that complements the ecosystem of infrastructures and open data about and for science, contributing to the possibility of change in evaluation models.
Understanding how data moves within the scientific process, how it contributes to other projects, and on this basis being able to generate metrics, criteria, and inform decisions, is fundamental for the advancement of science. These vital data sources will definitely contribute to the CRIS (Current Research Information Systems) that countries or institutions develop and implement.
Ricardo Hartley Belmar:
I am interested in the Data Citation Corpus because, while I currently extract data from the DataCite API to analyze datasets, the Corpus offers complementary information that enhances my analysis. The DataCite API provides detailed metadata about datasets, including persistent identifiers and relationships between resources. However, the Data Citation Corpus aggregates a centralized and publicly available resource of data citations from various sources, including mentions in articles and preprints, allowing for a deeper understanding of how datasets are cited and reused across different disciplines. In that way, I can generate valuable evidence on data citation and reuse practices, supporting efforts to promote research data’s open availability and reuse. This analysis is crucial for fostering a culture that recognizes data as an essential research output, vital for academic and societal progress.
Chokri Ben Romdhane:
First of all, the Corpus is an excellent resource for us to harvest data and enrich our repository with data citations using the Corpus data file – and later an API service, once developed. The Corpus is also a model for us to handle data citations in our repository, the Tunisian Technical and Scientific Information Portal, since it uses trusted infrastructure and consistent and transparent data models to incorporate data citations. By participating in the development of the Corpus, we will gain the necessary capacities to access and display citations to Tunisian data in our repository.
Barbora Bieliková:
Currently, mainly as a means of open science advocacy and an example of good practices. However, in the future, I hope to use the Corpus to demonstrate how our faculty´s researchers’ datasets are used and cited.
Why should institutions and funders pay attention to and use the Data Citation Corpus as it develops?
Francisco Silva-Garcés:
Open infrastructure, like the Data Citation Corpus, are built by the community. As is often the case with community-driven projects, it undergoes a process of evolution and continuous improvement with broad participation. What gives strength and vitality to open-source projects (whether software or not) is their community. In fact, the community is the thermometer of an open-source project. Institutions and funding agencies that do not keep track of the development of these projects and their communities are missing out on a potential opportunity to be part of something significant.
Institutions and funding agencies should also contribute to and integrate into these communities and learn from what is being developed, because they will inevitably need to incorporate new metrics and parameters into their policies.
Ricardo Hartley Belmar:
Institutions, like universities and funders, should engage with the evolving Data Citation Corpus because it offers a comprehensive view of how research data is utilized and reused, directly aligning with universities’ third mission of societal engagement. Incorporating this resource into evaluation processes could also provide more elements to understanding research impact (the “how”) and evaluate research outputs in a more holistic and informed way, even informing strategic decisions on resource allocation and future investments in research infrastructure. Additionally, the Corpus could be useful in generating insights into data reuse patterns, enabling organizations to assess the effectiveness of their data-sharing policies and identify datasets that influence the scientific community.
Chokri Ben Romdhane:
The Data Citation Corpus will be an open solution alternative for Tunisian researchers to find data in addition to the commercial tools. The Corpus will contribute to the improvement of the visibility of the Tunisian data locally and globally, which can in turn increase the reuse of those data. The Corpus will serve also as a tool for Tunisian research institutions to evaluate their data, and assess them within the regional and global context.
Barbora Bieliková:
I think that various institutions can benefit from the Data Citation Corpus in the same way as they leverage open tools to evaluate publication outputs – it is one of the ways of assessing the impact of the researchers’ work.