Kaggle competition

Make Data Count aims to support the community in understanding what data are used, when and how, so that we can establish and recognize the value and impact of open scientific data. To this end, Make Data Count maintains the Data Citation Corpus, an aggregation of links between data and papers, but these links are currently incomplete: they only cover a fraction of the published literature, and they do not provide context on how the data were used.
To address this, Make Data Count launched a Kaggle competition calling for state of the art machine learning models (including LLMs) to identify mentions of data in papers AND contextualize the relationship.
The competition will award prizes to highly-performant models that can identify mentions to data in the scientific literature, and classify the data citations as either primary (data generated as part of the paper) or secondary (data reused or derived from elsewhere).
The models are open source and will support work to automate text mining of the scientific literature to contribute high quality and contextualized data-to-paper connections to the MDC Data Citation Corpus.
Total prizes: $100,000
The competition has now closed.
We are grateful to Kaggle, The Navigation Fund, and Chan Zuckerberg Initiative for their support and are eager to use the winning models to advance the Data Citation Corpus for all.
Read more about the competition here