close icon

Data Citation Corpus documentation

The Data Citation Corpus is a project by DataCite and Make Data Count funded by the Wellcome Trust, which has as focus the development of a comprehensive, centralized and publicly-available resource of data citations from a variety of sources.

Current release

The fourth release of the Data Citation Corpus was delivered on July 27, 2025. This follows the first release of the Data Citation Corpus in January 2024, the second release on August 23, 2024 and the third release on February 1, 2025. The current release incorporates the following additions and improvements:

Additions

  • 5.2 million data citations ingested from Europe PMC on 9 July 2025, which employs text-mining algorithms to identify citations to datasets from over 40 life-sciences repositories in the main text and the supplementary files of articles indexed in Europe PMC. These data-article connections are available via the Europe PMC annotations service. Citations with provenance from Europe PMC are identified as “eupmc” in the source field. 
  • 139,647 additional data citations from Event Data for the period 1 January 2025 through 30 June 2025.

Metadata enhancements

  • Affiliation information retrieved from the GEO repository for citations to datasets from the Gene Expression Omnibus (GEO), as well as reconciliation to Research Organization Registry (ROR) where possible.
  • Reconciliation of organization and funder names with the Research Organization Registry (ROR) for new citations from Event Data.
  • Subject mapping for citations from Europe PMC corresponding to datasets with accession numbers.

The current data file for the Data Citation Corpus consists of 10.7 million records and a dashboard for visualizing the contents of the data file.

Data citation records

Each data citation record is comprised of a pair of identifiers i.e. an identifier for the dataset (a DOI or an accession number) and the DOI of the publication object (journal article or preprint) in which the dataset is cited, and various metadata for the dataset and for the citing object. An example data citation record is provided below:

{
“id”:”cdfe1599-7288-4fbd-b8d8-4a25fb6d170d”,
“created”:”2023-09-26T10:26:21.549Z”,
“updated”:”2023-11-15T23:17:25.755Z”,
“repository”:”European Nucleotide Archive”,
“publisher”:”F1000 Research Ltd”,
“journal”:”Wellcome Open Research”,
“title”:null,
“publication”:”10.12688/wellcomeopenres.17706.1″,
“dataset”:”OU744360.1″,
“publishedDate”:”2022-02-15T00:00:00.000Z”,
“source”:”czi”,
“subjects”:[],
“affiliations”:[],
“funders”:[]
}

Data file

The data file is available on Zenodo, in JSON and CSV formats: https://zenodo.org/records/11216814. The JSON file is the version of record.

Version 1.0 of the corpus data file was released on January 30, 2024. Release v1.1 provided an optimized version of v1.0 designed to make the original citation records more usable, with no change to the citations included. Release 2.0 reflected the removal of citations that were not in scope, the inclusion of new data citations created in DataCite Event Data since the first release and inclusion of disciplinary information for citations for data with accession numbers. Release 3.0 incorporated data citations contributed by the Aligning Science Across Parkinson’s (ASAP) initiative, new citations in DataCite Event Data since the second release and addition of ROR information for affiliation and funder metadata. Release 4.0 includes data citations ingested from Europe PMC, new citations in DataCite Event Data in the first six months of 2025, and addition of affiliation metadata for citation records for datasets from the Gene Expression Omnibus (GEO) repository. Release 4.0 also includes disciplinary information mapping for citations for data with accession numbers, and ROR matching for citations from DataCite Event Data and for datasets from GEO.

Feedback on the data file can be submitted via Github. For general questions, email info@makedatacount.org.

Data Sources

The file includes four sources for data citations:

DataCite Event Data: Citations are determined based on the resource type and the relation type designated in the metadata for the dataset or the article (see DataCite documentation on contributing citations):

ResourceType= Dataset; relationType=IsReferencedBy/IsCitedBy/IsSupplementTo
ResourceType= Text; relationType= References/Cites/IsSupplementedBy

Event Data includes some citations originating from article metadata registered at Crossref, these may carry relation types different from the three listed above for DataCite metadata.

Chan Zuckerberg (CZI) Science Knowledge Graph: Mention to dataset identifier (accession number or DOI) identified in the text of an article by NER Model (SciBERT Model)

Aligning Science Across Parkinson’s (ASAP): Mention to dataset identifier (accession number or DOI) identified in the text of an article by DataSeer through machine learning software, or via a manuscript’s Key Resource Table, and verified by a DataSeer and/or an ASAP curator.

Europe PMC: Mention to dataset identifier (accession number or DOI) identified via text mining in the text or supplemental files of an article indexed in Europe PMC. For the Europe PMC approach, see https://gitlab.ebi.ac.uk/literature-services/public-projects/textmining-dictionaries/. For the Data Citation Corpus, we ingested data citations available via Europe PMC’s annotations files, including those in scope for the Data Citation Corpus, per the summary in the table below.

 

Platforms in Europe PMC filesIn scope for Data Citation Corpus?
alphafoldYes
arrayexpressYes
bia (BioImage Archive)Yes
biomodelsYes
bioprojectYes
biosampleYes
biostudiesYes
BRENDANo
cathYes
cellosaurusNo
chebiNo
chemblYes
complexportalYes
dbgapYes
doiYes
ebiscNo
efo (Experimental Factor Ontology)No
emdbYes
egaYes
empiarYes
ensemblYes
eudractNo
gcaYes
genYes
geoYes
gisaidYes
goNo
hgncYes
hipsciNo
hpaNo
igsrNo
intactYes
interproYes
metabolightsYes
metagenomicsYes
mintNo
nctNo
omimNo
orphadataYes
pdbYes
pfamYes
pxdYes
reactomeYes
refseqYes
refsnpYes
rfamNo
rheaNo
rnacentralYes
rridNo
treefamYes
uniparcYes
uniprotYes

 

Note that each citation record designates a data citation by one source. There is a proportion of data citations which have been asserted by more than one source, i.e. some citation records designate the same dataset ID and article DOI pair, but a different source. These instances are currently listed as separate citation records. The Data Citation Corpus includes 9.6 million unique records (i.e. citation records with the same dataset ID-article DOI pair).

The scope of the data file covers dataset-article pairs, i.e. it includes pairs where the citing object is a journal article or a preprint and the cited object is a dataset. DataCite Event Data includes records where the citing object involves a range of resource types (e.g. datasets, software), for the purposes of the data file of the Corpus, the only citations from DataCite Event Data included are those where the cited object is a dataset and the citing object is an article.

In addition to the identifier for the dataset and the citing object, each record includes metadata fields for the journal, publisher and publication date for the citing object (from Crossref metadata) and the repository where the dataset is hosted (via DataCite or EMBL-EBI). Where additional metadata fields are available (e.g. for affiliation, subject or other) this is included in the data citation record. Coverage of these additional metadata fields varies across citations.

Data Structure

Each data citation record includes the following fields:

FieldDescriptionRequired?
idInternal identifier for the citationYes
createdDate of item’s incorporation into the corpusYes
updatedDate of item’s most recent update in corpusYes
repositoryRepository where cited data is storedNo
publisherPublisher for the article citing the dataNo
journalJournal for the article citing the dataNo
titleTitle of cited dataNo
publicationDOI of article where data is citedYes
datasetDOI or accession number of cited dataYes
publishedDateDate when citing article was publishedNo
sourceSource where citation was harvestedYes
subjectsSubject information for cited dataNo
affiliationsAffiliation information for creator of cited dataNo
fundersFunding information for cited dataNo

 

Dashboard

The dashboard at https://corpus.datacite.org/dashboard provides an overview of the content of the data file. This includes the six visualizations below with filtering options according to different facets (e.g. affiliation, repository, journal etc):

  • Citation counts over time: Count of data citations spanning the time frame of currently available corpus data, visualized from 2015 to 2025.
  • Citation counts by publisher: Count of data citations by publisher.
  • Counts of unique repositories, journals, subjects, affiliations, funders: Breakdown of the current coverage in the corpus for journals, affiliations, repositories and subjects.
  • Citation counts by subject: Count of data citations per dataset subject, note this displays the distribution of records that contain this metadata field and is not representative of the full set of records in the data file.
  • Citation counts by source of citation: Counts for citations ingested from DataCite Event Data, CZI Science Knowledge Graph, ASAP and Europe PMC.
  • Data citations corpus growth: Citation counts and ingest date (into the corpus) by identifier type over time.

Existing Limitations & Planned Enhancements

The Data Citation Corpus brings together for the first time data citations associated with datasets with DOIs and accession number IDs. There are a few limitations and considerations that users should bear in mind when analyzing data in the current data file. We will continue our work to address these limitations in the course of future development of the Corpus.

Metadata coverage

As noted above, coverage of metadata fields varies across records. The current release incorporates disciplinary information mapping for datasets with accession numbers, and provides affiliation metadata for datasets from the GEO repository. We have performed ROR reconciliation for GEO dataset records and citations from Event Data that have affiliation strings and we included ROR IDs and ROR names for matches that return a value of chosen: TRUE per the ROR matching algorithm. We will work to add affiliation metadata for additional data citations, along with ROR IDs where possible.

NER model output

The data citations contributed by Chan Zuckerberg Initiative were identified by a NER model – the methodology for this is outlined in the Appendix below. We have undertaken steps to minimize the number of false positives identified through the NER model by completing checks against the expected accession number structure for datasets coming from repositories with accession numbers. This resulted in the removal of 473,792 data citations previously included in the Corpus. We are confident this has increased the accuracy of data citations in the Corpus, but note to users the possibility of some additional false positives remaining in the data file.

Disciplinary coverage

The data citations identified via CZI’s NER model and by Europe PMC involve mining articles indexed in Europe PMC, which has a biomedical scope. The accession numbers are associated with repositories also focused on life sciences. As a result, the data citations from these two sources originate mostly from disciplines in the life sciences. We will seek to extend the disciplinary coverage of the Data Citation Corpus as we ingest data citations from additional sources.

Appendix: CZI Science methodology

Full-text articles included

The set of articles employed for text mining involved 5.3 million articles, where the full text was available open access in Europe PMC.

Repositories mined

The list of repositories the NER model mined for are listed below, with a row entry for DOIs.

  • List of terms mined for come from https://europepmc.org/pub/databases/pmc/TextMinedTerms/
  • All but three repositories were linked through identifiers.org, those not linked are ebisc, gisaid, hipsci 
  • For the purposes of the data file, mentions to identifiers to eudract and nct were excluded as those are clinical trial registries. Per trial best practices, clinical trials are generally registered prior to the recruitment of patients, as a result, there is no guarantee that the clinical trial record will include a dataset from the trial.

 

Repository nameIdentifier prefixLinking Methodology

Where data is the extracted_word (or data mention)

Linked through identifiers.org

Is the link an identifiers.org link?

ArrayExpressarrayexpresshttps://identifiers.org/arrayexpress:datasetY
BioModelsbiomodels.dbhttps://identifiers.org/biomodels.db:datasetY
BioProjectbioprojecthttps://identifiers.org/bioproject:datasetY
biosamplehttps://identifiers.org/biosample:datasetY
BioStudiesbiostudieshttps://identifiers.org/biostudies:datasetY
CATHcathhttps://identifiers.org/cath:datasetY
chebihttps://identifiers.org/chebi:datasetY
ChEMBLchemblhttps://identifiers.org/chembl:datasetY
Complex Portal (CP)complexportalhttps://identifiers.org/complexportal:datasetY
NCBI dbGaP (DataBase of Genotypes And Phenotypes)dbgaphttps://identifiers.org/dbgap:datasetY
doihttps://dx.doi.org/:datasetsometimes
EBiSC Catalogue (European Bank for induced pluripotent Stem Cells catalogue)ebischttps://cells.ebisc.org/datasetN
Experimental Factor Ontologyefohttps://identifiers.org/efo:datasetY
The European Genome-phenome Archive(EGA)egahttps://identifiers.org/ega.dataset:datasetY
The Electron Microscopy Data Bank (EMDB)emdbhttps://identifiers.org/emdb:datasetY
Electron Microscopy Public Image Archive (EMPIAR)empiarhttps://identifiers.org/empiar:datasetY
Ensemblensemblhttps://identifiers.org/ensembl:datasetY
EU Clinical Trial Register(EUCTR)eudracthttps://identifiers.org/euclinicaltrials:datasetY
Genome assembly databasegcahttps://identifiers.org/insdc.gca:datasetY
European Nucleotide Archivegenhttps://identifiers.org/ena.embl:datasetY
Gene Expression Omnibus (GEO)geohttps://identifiers.org/geo:datasetY
GISAIDgisaidhttp://gisaid.org/EPI/datasetN
Gene Ontologygohttps://identifiers.org/go:datasetY
HUGO Gene Nomenclature Committeehgnchttps://identifiers.org/hgnc:datasetY
Human induced pluripotent stem cell initiativehipscihttp://www.hipsci.org/lines/#/lines/datasetN
The Human Protein Atlashpahttps://identifiers.org/hpa:datasetY
The International Genome Sample Resourceigsrhttps://identifiers.org/coriell:datasetY
IntActintacthttps://identifiers.org/intact:datasetY
InterProinterprohttps://identifiers.org/interpro:datasetY
MetaboLightsmetabolightshttps://identifiers.org/metabolights:datasetY
MGnifymetagenomicshttps://identifiers.org/mgnify.samp:datasetY
minthttps://identifiers.org/mint:datasetY
ClinicalTrials.govncthttps://identifiers.org/clinicaltrials:datasetY
omimhttps://identifiers.org/mim:datasetY
Orphadataorphadatahttps://identifiers.org/orphanet:datasetY
The Protein Data Bankpdbhttps://identifiers.org/pdb:datasetY
Pfam Protein Familiespfamhttps://identifiers.org/pfam:datasetY
PRIDE Proteomics Identification Database*pxdhttps://identifiers.org/pride:datasetY
Reactomereactomehttps://identifiers.org/reactome:datasetY
NCBI Reference Sequence Databaserefseqhttps://identifiers.org/refseq:datasetY
dbSNP Reference SNPrefsnphttps://identifiers.org/dbsnp:datasetY
Rfamrfamhttps://identifiers.org/rfam:datasetY
RNAcentralrnacentralhttps://identifiers.org/rnacentral:datasetY
Research Resource Identifiersrridhttps://identifiers.org/rrid:datasetY
TreeFamtreefamhttps://identifiers.org/treefam:datasetY
uniparchttps://identifiers.org/uniparc:datasetY
UniProtuniprothttps://identifiers.org/uniprot:datasetY

*Listed as deactivated in identifiers.org as of 22 Mar 2024

DOIs

DOIs are particularly noisy as the model does not always distinguish between data and non-data DOIs. The primary reason being that there are many article-article citations in reference lists and the model could not easily distinguish this. In order to try and address this, we excluded reference lists from the text mining of articles, and completed an additional content negotiation step to verify that the DOI corresponds to a dataset. We intend to work with the community to better address this in future iterations.