Applying nuance to our understanding of data metrics for evaluation
March 5, 2025 | By: Clare DeanPerspectives from Nicolás Robinson-García

Nicolás Robinson-García works in bibliometrics and research evaluation. His current research focuses on research contributions and how those are distributed across teams and in different research fields – see this preprint for his latest research on contributions based on the CRediT taxonomy.
Nicolás has previously dived into researchers’ attitudes about their diverse outputs including data, and their place in evaluation systems. His work has involved exploring different data sources for bibliometric analyses and their metadata coverage (see their analysis of the Data Citation Index and DataCite), including connections between datasets and articles. Given his breadth of expertise in research activities and evaluation, we asked Nicolás for his perspectives on the challenges we’re facing in the development of responsible metrics for research data.
Nicolás highlighted the importance of context as part of the evaluation of individual researchers. For any given researcher profile, we have to understand the context for the field where the research is carried out, and what tasks and activities that research involves. The outputs a researcher produces may also vary from one discipline to another. This context is necessary in order to see whether and how those aspects need to be considered as part of the evaluation process, i.e. how do we benchmark and normalize activities within specific roles and fields. Responsible metrics for data, therefore, must account for variations in researcher profiles and activities. This also means that we need a nuanced view into the indicators of data usage and impact, in some situations or disciplines, data downloads or views may be more important than other indicators such as citations.
Beyond disciplinary differences, there are additional dimensions that are relevant but can bring challenges. Datasets can have different versions so it is critical to understand what the point of entry is for the user of the dataset, what version is being used and cited. This is important for reproducibility and for accuracy in evaluation. However, for a number of data types, establishing the version can be challenging, for example, when interacting with a live dataset.
A step forward in addressing these aspects is to utilize persistent identifiers, and to leverage their metadata to establish connections between contributors and the dataset; as we move ahead, such connections would ideally be enriched with additional context about what uses the researchers made of the data.
Nicolás has used a number of research information resources in his research, including DataCite, Web of Science, ORCID and others. We asked for his views on what are important considerations for the development and characteristics of these resources, to make them useful as a source for meta-research studies. Information accuracy and cleanliness came up as immediate considerations. It is common for databases for research information to require clean up and troubleshooting of errors to narrow down the records relevant to the research question at hand. The cleaner the data in the resource, the more useful it will be for meta-researchers.
When looking at data specifically, an important aspect is how the connection is established to the article. There can be a number of biases that go into what constitutes a ‘data’ record and how the connection is established, which it’s important to be mindful of. But once these article-data connections are established, they can open a number of avenues for research, including: what is the discovery path for the dataset and is there an association with its reuse? What does this look like for interdisciplinary research projects and teams? What information can we gain about how research teams work together across the project cycle?
Nicolás maintained that in order to facilitate more meta-research into data practices and data evaluation, information about data contributions and the connections to these should be openly available. This can relate to data-to-article connections (through citations or data availability statements) but also other links such as data-software connections. A significant aspect of the metadata that will be important is affiliation metadata. Often this is not openly available or difficult to access, and having more affiliation metadata openly available will enable more meta-research studies.
| Would you like to share your views on data evaluation and assessment? Contact us about sharing your perspective and experience on our blog. |