Perspectives on data evaluation and assessment
February 14, 2025 | By: Make Data CountViews from Peter Cerda, Data Curation Specialist
Peter Cerda is a Data Curation Specialist for Workflows and Big Data with the library at the University of Michigan, Ann Arbor. He helps students and faculty deposit large, complex datasets into Deep Blue Data Repository; and helps researchers think about how they can best manage and preserve their data to make it more accessible and reusable to the research community and the public. We asked Peter about his views on the importance of data evaluation and the progress of infrastructure to support it.
Why do you believe evaluation of the use and reach of open data is important?
Researchers spend a lot of time and funds to collect and process data for their own research. Sharing this data openly allows others, including those who may have limited resources, to use it. I also believe that researchers should get more credit for their work. If they’re making really important datasets available, they should be able to receive credit related to their use, to how those datasets are contributing to further research activities. It’s vital to have metrics that showcase how open datasets are viewed and used by others, and which can provide insights into how researchers’ data contributions are tied to important developments within their fields.
What are examples of progress in data evaluation you would like to highlight?
Relative to our data repository at the University of Michigan, we’ve been encouraging people to cite their data in their articles and other outputs. We integrated ORCID into our depositing platform so that datasets can be linked to researchers’ ORCIDs and gain wider visibility.
We’re part of the Data Curation Network (DCN). My colleagues at DCN advocate for open data, and have been doing some research to evaluate how data is being used. We’re looking into what purposes people are actually using their data for. We’ve put a lot of time and investment into open data, but need to evaluate how people are using it. One of the things we’re trying to implement at our institution is a brief survey for users of the repository, asking visitors how they found the dataset, what are their intended uses for it, etc. Ultimately, we’d also like to ask about their affiliation, and role i.e. whether they are a member of the public, a government representative, working in the private sector etc. This kind of information is important in building a picture of the use of open data.
Why are you interested in Make Data Count’s Data Citation Corpus? How have you engaged with it?
We’re interested in uncovering trends for our researchers, exploring how their datasets are being cited, including self-citation. We want to know how people are downloading and citing our datasets. We looked at the Corpus data file, and isolated our individual records to explore those connections. As the Corpus continues to develop, it will help us gain insights on citation trends overall and for our institution, and to potentially explore how interactions with our data compare to those from other repositories.
What do you think the value of the Data Citation Corpus is, or could be to researchers?
I think that the ultimate value for researchers is that the Data Citation Corpus could be the place to go to show the value of your data. Many researchers use Google Scholar as the platform to see citations to their papers. The Data Citation Corpus could contribute something of similar value – a single dashboard that a researcher can visit that shows the connections to and value of their data. There are a number of platforms you can go to find information about research performance but they focus on articles, the Data Citation Corpus would provide a single place to go for data citation.
What potential developments would you like to see in the Data Citation Corpus in alignment with your institution’s interests?
I’d like to see more data citations added over time and more metadata, such as institutional affiliations, ORCID, and more metadata fields to enrich the output and usefulness of the Corpus.
What do you think needs to change to advance data evaluation and recognition as part of institutions’ research assessment processes?
It’s important to get buy-in from the institutions to recognize that data metrics and evaluation matter, and that data are a legitimate scholarly output. Data evaluation should be considered in tenure and promotion processes. Funders’ focus has so far been on the sharing of data, but it would be useful if they also expressed interest in data assessment – downloads, views, etc. There seems to be more interest in data evaluation and recognition from early career researchers, and until now there hasn’t been an easy way to show the number of citations they’ve received. Combined top down and bottom up approaches will be important to generate support for and progress data evaluation.
| Would you like to share your views on data evaluation and assessment? Contact us about sharing your perspective and experience on our blog. |