close icon

Opening Up Translational Science Data Impact Through the Data Citation Corpus

May 13, 2025   |   By: Make Data Count

Reese Richardson (Engineering Sciences and Applied Mathematics, Northwestern University), Iratxe Puebla (Make Data Count), Jason Portenoy (DataCite), Karen Gutzman (Galter Health Sciences Library, Northwestern University Feinberg School of Medicine) and Kristi Holmes (Galter Health Sciences Library and Northwestern University Clinical and Translational Sciences Institute, Northwestern University Feinberg School of Medicine)

DOI: 10.60804/27r9-2g67

As open science practices continue to gain momentum, understanding how data itself advances science and benefits patients is more important than ever. Open datasets can catalyze new discoveries and enable collaborations across disciplines, supporting progress in clinical care – but can we actually measure that impact?

Through a recent collaboration between researchers at Northwestern University and Make Data Count, we carried out an exploratory analysis of the Data Citation Corpus to begin to probe how biomedical datasets are used and reused, and what that reuse tells us about their translational impact: how does knowledge move through datasets from basic science into real-world health applications.

Identifying data reuse with the Data Citation Corpus

We were able to use the metadata from the five million available data citations in the Data Citation Corpus as the source for the analysis. For this exploratory study, we honed in on a specific slice of biomedical data: datasets in the Gene Expression Omnibus (GEO), which archives microarray, next-generation sequencing, and high-throughput functional -omics data. We focused on GEO datasets because this repository assigns structured identifiers, allowing greater confidence that the set of data citations in the Corpus identified via text-mining methodologies correlate with a data identifier (see the Data Citation Corpus documentation for information about the text-mining approach for data citations). Using GEO metadata, we identified microarray datasets in human cells and tissues initially published from 2005 to 2009 (n=3,427) and looked for citations to these datasets in the Corpus. We then used iCite to find articles that cited these datasets that were themselves cited by clinical trials.

Our exploratory analysis suggested that larger datasets (i.e., with a greater number of biological samples) were more likely to be highly-cited datasets and were more likely to be used repeatedly in articles that were cited by clinical trials (Figure 1).

Figure 1. Scatter plots of sample counts vs. number of citations for GEO datasets in the Data Citation Corpus. Dataset citations were collected through the Data Citation Corpus v3.0, inter-article citations were collected through iCite v32.

Tracing the translational impact of datasets

Two highly-cited datasets caught our attention:

These datasets, one from a follow-up study on the other, were the highest-cited among articles that were subsequently studied in clinical trials. These datasets and the articles reporting on them appeared to have been hugely influential in establishing the practice of using -omics data to develop “prognostic signatures”. 

We examined the citation networks between the GEO datasets and the original articles reporting these two datasets, as well as the publications that subsequently cited those articles. This analysis identified 55 non-clinical trial articles that cite the GSE2034 and GSE7390 datasets that were later cited by 124 clinical trials (Figure 2).

 

Figure 2. Network diagram of data citations in the Data Citation Corpus for the GEO datasets GSE2034 and GSE7390. The GEO datasets are shown in green. The original articles reporting on these datasets are shown in orange. The 55 non-clinical trial articles that cite these datasets and were later cited by clinical trials are shown in blue. The 124 clinical trials that subsequently cite these articles are shown in red. Citations between these articles and datasets are shown with arrows directed from citing work to referenced work. 

Open data fuels translational science

This kind of analysis highlights the translational value of open data. It also shows how the analysis of clinical trial research can provide a valuable indicator of the health relevance of data contributions. Finally, it showcases the dynamic life cycle of open data: datasets are dynamic resources that support scientific endeavours long after they become available. In this case, the impact of these two datasets has persisted almost two decades since their publication. 

While here we focused on citation networks for GEO datasets, there are additional opportunities for analyzing dataset impact by exploring datasets associated with specific affiliations, funders or research fields. Such facets provide opportunities for metadata enrichment and they are areas we are actively working on for the Data Citation Corpus. We hope to also analyze the roles of metadata quality, citation practices, and repository standards in making data truly reusable. 

Our analysis showcases examples of valuable datasets in translational research. Today, we have tools that allow us to dive deeper into the use of datasets, and we hope that the community will pursue further exploration of the reach and impact of open datasets. By understanding how datasets move through the research ecosystem and foster new innovations, we can build stronger incentives for open, reusable data practices and recognition for data as a primary research output.

Read the poster ‘Opening up translational data impact through the Data Citation Corpus’ on Zenodo