close icon

Data Citation Corpus documentation release 2.0

The Data Citation Corpus is a project by DataCite and Make Data Count funded by the Wellcome Trust, which has as focus the development of a comprehensive, centralized and publicly-available resource of data citations from a variety of sources.

Current release

The first release of the Data Citation Corpus was delivered on January 30, 2024. The second release was shared on August 23, 2024 and incorporates the following additions and improvements compared to the initial release:

Additions

  • 179,885 additional data citations from Event Data for the period 1 June 2023 through 30 June 2024
  • Inclusion of Field of Science subject terms for data citation records originating from CZI, based on the disciplinary area(s) of the data repository

Data citation record improvements

  • Removal of data citations deemed out of scope for the Data Citation Corpus: citations to repositories for outputs other than datasets, records where the metadata relationship types or the dataset/article pairs did not fit the criteria for a data citation
  • Removal of 473,792 records from citations originating from CZI, where the accession number did not match the expected structure for identifiers at the designated repository
  • Removal of 4,110,019 records from citations originating from CZI, involving duplication of data-article pairs, for these instances the record with the most recent updated date was retained
  • Removal of personally identifying information from affiliation and funder metadata fields

The current data file for the Data Citation Corpus consists of 5 million data citations and a dashboard for visualizing the contents of the data file.

Information about the Data Citation Corpus is also available via DataCite documentation

Data citation records

Each data citation record is comprised of a pair of identifiers i.e. an identifier for the dataset (a DOI or an accession number) and the DOI of the publication object (journal article or preprint) in which the dataset is cited, and various metadata for the dataset and for the citing object.

{
“id”:”cdfe1599-7288-4fbd-b8d8-4a25fb6d170d”,
“created”:”2023-09-26T10:26:21.549Z”,
“updated”:”2023-11-15T23:17:25.755Z”,
“repository”:”European Nucleotide Archive”,
“publisher”:”F1000 Research Ltd”,
“journal”:”Wellcome Open Research”,
“title”:null,
“publication”:”10.12688/wellcomeopenres.17706.1″,
“dataset”:”OU744360.1″,
“publishedDate”:”2022-02-15T00:00:00.000Z”,
“source”:”czi”,
“subjects”:[],
“affiliations”:[],
“funders”:[]
}

Data file

The data file is available on Zenodo, in JSON and CSV formats: https://zenodo.org/records/11216814. The JSON file is the version of record.

Version 1.0 of the corpus data file was released on January 30, 2024. Release v1.1 provided an optimized version of v1.0 designed to make the original citation records more usable, with no change to the citations included. Release 2.0 reflects the removal of citations that were not in scope, or have been determined to be false positives, and the inclusion of new data citations created in DataCite Event Data since the first release. In addition, release 2.0 incorporates metadata enhancements to list disciplinary information for citations for data with accession numbers, and increased accuracy to the affiliation and funder information, where this is available.

Feedback on the data file can be submitted via Github. For general questions, email info@makedatacount.org.

Data Sources

The file includes two sources for data citations:

DataCite Event Data: Citations are determined based on the resource type and the relation type designated in the metadata for the dataset or the article (see DataCite documentation on contributing citations):

ResourceType= Dataset; relationType=IsReferencedBy/IsCitedBy/IsSupplementTo
ResourceType= Text; relationType= References/Cites/IsSupplementedBy

Event Data includes some citations originating from article metadata registered at Crossref, these may carry relation types different from the three listed above for DataCite metadata.

Chan Zuckerberg (CZI) Science Knowledge Graph: Mention to dataset identifier (accession number or DOI) identified in the text of an article by NER Model (SciBERT Model)
More information about open sourced algorithm is forthcoming.

The scope of the data file covers dataset-article pairs, i.e. it includes pairs where the citing object is a journal article or a preprint and the cited object is a dataset. DataCite Event Data includes records where the citing object involves a range of resource types (e.g. datasets, software), for the purposes of the data file of the Corpus, the only citations from DataCite Event Data included are those where the cited object is a dataset and the citing object is an article.

In addition to the identifier for the dataset and the citing object, each record includes metadata fields for the journal, publisher and publication date for the citing object (from Crossref metadata) and the repository where the dataset is hosted (via DataCite or EMBL-EBI). Where additional metadata fields are available (e.g. for affiliation, subject or other) this is included in the data citation record. Coverage of these additional metadata fields varies across citations.

Data Structure

Each data citation record includes the following fields:

Dashboard

The dashboard at https://corpus.datacite.org/dashboard provides an overview of the content of the data file. This includes the six visualizations below with filtering options according to different facets (e.g. affiliation, repository, journal etc):

  • Citation counts over time: Count of data citations spanning the time frame of currently available corpus data, from 2013 to 2023.
  • Citation counts by publisher: Count of data citations by publisher.
  • Counts of unique repositories, journals, subjects, affiliations, funders: Breakdown of the current coverage in the corpus for journals, affiliations, repositories and subjects.
  • Citation counts by subject: Count of data citations per dataset subject, note this displays the distribution of records that contain this metadata field and is not representative of the full set of records in the data file.
  • Citation counts by source of citation: Counts for citations ingested from DataCite Event Data and CZI Science Knowledge Graph.
  • Data citations corpus growth: Citation counts and ingest date (into the corpus) by identifier type over time.

Existing Limitations & Planned Enhancements

The Data Citation Corpus brings together for the first time data citations associated with datasets with DOIs and accession number IDs. There are a few limitations and considerations that users should bear in mind when analyzing data in the current data file. We will continue our work to address these limitations in the course of future development of the Corpus.

Metadata coverage

As noted above, coverage of metadata fields varies across records. The current release incorporates subject information for a considerable number of the citations in the Corpus. We will work to add additional information to the existing data citations regarding affiliation details (ROR IDs and disambiguation of affiliation information) as well as funder details with ROR ID or Crossref Funder ID.

NER model output

The data citations contributed by Chan Zuckerberg Initiative were identified by a NER model – the methodology for this is outlined in the Appendix below. We have undertaken steps to minimize the number of false positives identified through the NER model by completing checks against the expected accession number structure for datasets coming from repositories with accession numbers. This resulted in the removal of 473,792 data citations previously included in the Corpus. We are confident this has increased the accuracy of data citations in the Corpus, but note to users the possibility of some additional false positives remaining in the data file.

Disciplinary coverage

The data citations identified via CZI’s NER model involved mining a set of articles indexed in Europe PMC, which has a biomedical scope. The accession numbers were associated with repositories also focused on life sciences disciplines. As a result, the data citations identified will be originating mostly from disciplines in the life sciences. The repositories included are well established in their fields and attract wide use in those disciplines, and so they provided a good starting point to identify citations to accession number IDs, but we will seek to extend the disciplinary coverage of the Data Citation Corpus as we ingest data citations from additional sources.

Appendix: CZI Science methodology

Full-text articles included

The set of articles employed for text mining involved 5.3 million articles, where the full text was available open access in Europe PMC.

Repositories mined

The list of repositories the NER model mined for are listed below, with a row entry for DOIs.

  • List of terms mined for come from https://europepmc.org/pub/databases/pmc/TextMinedTerms/
  • All but three repositories were linked through identifiers.org, those not linked are ebisc, gisaid, hipsci 
  • For the purposes of the data file, mentions to identifiers to eudract and nct were excluded as those are clinical trial registries. Per trial best practices, clinical trials are generally registered prior to the recruitment of patients, as a result, there is no guarantee that the clinical trial record will include a dataset from the trial.

 

Repository nameIdentifier prefixLinking Methodology

Where data is the extracted_word (or data mention)

Linked through identifiers.org

Is the link an identifiers.org link?

ArrayExpressarrayexpresshttps://identifiers.org/arrayexpress:datasetY
BioModelsbiomodels.dbhttps://identifiers.org/biomodels.db:datasetY
BioProjectbioprojecthttps://identifiers.org/bioproject:datasetY
biosamplehttps://identifiers.org/biosample:datasetY
BioStudiesbiostudieshttps://identifiers.org/biostudies:datasetY
CATHcathhttps://identifiers.org/cath:datasetY
chebihttps://identifiers.org/chebi:datasetY
ChEMBLchemblhttps://identifiers.org/chembl:datasetY
Complex Portal (CP)complexportalhttps://identifiers.org/complexportal:datasetY
NCBI dbGaP (DataBase of Genotypes And Phenotypes)dbgaphttps://identifiers.org/dbgap:datasetY
doihttps://dx.doi.org/:datasetsometimes
EBiSC Catalogue (European Bank for induced pluripotent Stem Cells catalogue)ebischttps://cells.ebisc.org/datasetN
Experimental Factor Ontologyefohttps://identifiers.org/efo:datasetY
The European Genome-phenome Archive(EGA)egahttps://identifiers.org/ega.dataset:datasetY
The Electron Microscopy Data Bank (EMDB)emdbhttps://identifiers.org/emdb:datasetY
Electron Microscopy Public Image Archive (EMPIAR)empiarhttps://identifiers.org/empiar:datasetY
Ensemblensemblhttps://identifiers.org/ensembl:datasetY
EU Clinical Trial Register(EUCTR)eudracthttps://identifiers.org/euclinicaltrials:datasetY
Genome assembly databasegcahttps://identifiers.org/insdc.gca:datasetY
European Nucleotide Archivegenhttps://identifiers.org/ena.embl:datasetY
Gene Expression Omnibus (GEO)geohttps://identifiers.org/geo:datasetY
GISAIDgisaidhttp://gisaid.org/EPI/datasetN
Gene Ontologygohttps://identifiers.org/go:datasetY
HUGO Gene Nomenclature Committeehgnchttps://identifiers.org/hgnc:datasetY
Human induced pluripotent stem cell initiativehipscihttp://www.hipsci.org/lines/#/lines/datasetN
The Human Protein Atlashpahttps://identifiers.org/hpa:datasetY
The International Genome Sample Resourceigsrhttps://identifiers.org/coriell:datasetY
IntActintacthttps://identifiers.org/intact:datasetY
InterProinterprohttps://identifiers.org/interpro:datasetY
MetaboLightsmetabolightshttps://identifiers.org/metabolights:datasetY
MGnifymetagenomicshttps://identifiers.org/mgnify.samp:datasetY
minthttps://identifiers.org/mint:datasetY
ClinicalTrials.govncthttps://identifiers.org/clinicaltrials:datasetY
omimhttps://identifiers.org/mim:datasetY
Orphadataorphadatahttps://identifiers.org/orphanet:datasetY
The Protein Data Bankpdbhttps://identifiers.org/pdb:datasetY
Pfam Protein Familiespfamhttps://identifiers.org/pfam:datasetY
PRIDE Proteomics Identification Database*pxdhttps://identifiers.org/pride:datasetY
Reactomereactomehttps://identifiers.org/reactome:datasetY
NCBI Reference Sequence Databaserefseqhttps://identifiers.org/refseq:datasetY
dbSNP Reference SNPrefsnphttps://identifiers.org/dbsnp:datasetY
Rfamrfamhttps://identifiers.org/rfam:datasetY
RNAcentralrnacentralhttps://identifiers.org/rnacentral:datasetY
Research Resource Identifiersrridhttps://identifiers.org/rrid:datasetY
TreeFamtreefamhttps://identifiers.org/treefam:datasetY
uniparchttps://identifiers.org/uniparc:datasetY
UniProtuniprothttps://identifiers.org/uniprot:datasetY

*Listed as deactivated in identifiers.org as of 22 Mar 2024

DOIs

DOIs are particularly noisy as the model does not always distinguish between data and non-data DOIs. The primary reason being that there are many article-article citations in reference lists and the model could not easily distinguish this. In order to try and address this, we excluded reference lists from the text mining of articles, and completed an additional content negotiation step to verify that the DOI corresponds to a dataset. We intend to work with the community to better address this in future iterations.