close icon

The journey of data: Meta-researcher perspectives on data use and metrics

April 14, 2025   |   By: Make Data Count

DOI 10.60804/ZGJX-0K93


Over the last couple of months we have talked with meta-researchers about their work on data, and their perspectives on steps needed to better understand how datasets are used and inform responsible metrics for data. These conversations have highlighted a strong interest in understanding how data practices are evolving, and in developing evidence about how data travels through and advances research activities. This is part of a wider motivation in meta-research to study open science, looking also into connections between data and software, and what open outputs can tell us about the dynamics of research collaborations. These insights are valuable to inform our activities at Make Data Count, and I share below my key takeaways from these conversations.

 

Data sharing and data use differ across disciplines

There are differences in data sharing and data use across research areas, and thus, we should be mindful of not pushing a common set of mandates and/or incentives across fields. Instead, we should prioritize obtaining discipline-level evidence so that we can account for these differences and benchmark and normalize practices for different fields. This also involves understanding what indicators of use (citations, views, downloads, qualitative information, or other) better reflect the research activities of different communities. 

Data use opens up avenues to better understand interactions throughout the research cycle

Datasets are different from journal articles. Data may take different representations (a record at a repository or at a knowledge base, a file associated with an article, and others), and are more easily versioned. This means that often users access and interact with datasets in more diverse ways than they do with journal publications. Datasets are also used and created at multiple stages of the research process, so they can bring more granular insights into interactions throughout that process compared to article citations.This opens up a range of new questions to explore: What is the discovery path for the dataset and how does this influence its reuse? What do dataset connections to different objects tell us about researcher contributions across a project? Who are the data creators and the data users, and how does this differ across disciplines? 

AI and data readiness are key considerations to support further meta-research

Meta-research on the use of data requires samples of an adequate size to allow meaningful analysis and insights. AI will play an increasingly important role toward this, as  machine-learning approaches can surface data-article connections at a scale not previously possible. The Data Citation Corpus is an example of ongoing work to scale the data citations available to the community. The project leverages the potential of AI methodologies to surface data-article connections beyond those captured in the bibliographies, providing a more comprehensive view into the use of data as part of publications. To increase the visibility of data usage indicators, and to ensure they can enrich our pool of research information, it will be important to integrate information about data objects and their use into bibliographic databases.

Beyond scale, usage information also needs to be as ‘analysis-ready’ as possible, to allow researchers to narrow it down to what is relevant to their research questions. For data usage indicators to be valuable for meta-research and to build the foundations for responsible data metrics, we will need to keep sight both on quantity and quality.

Advancing data recognition requires a combination of approaches

Both top-bottom and bottom-up approaches will be needed to bring better recognition for data contributions in research assessment. Policies that outline an expectation of data sharing and processes that provide dedicated space to report on data contributions will be important in order to signal to researchers that they can report datasets as a primary contribution in their own right. In parallel to this, examples of the reach and broad use of research data from different projects will be powerful to showcase that data impact is not an aspirational goal, but rather is already taking place across different disciplines.

Do you have perspectives about meta-research evidence needed to inform data metrics? Get in touch if you’d like to share your experiences and views.