close icon

Announcing Make Data Count’s Kaggle Competition

June 11, 2025   |   By: Make Data Count

Post by Make Data Count advisors Daniella Lowenberg and Jennifer Lin

DOI: 10.60804/asfb-f691

We are thrilled to announce the launch of our Kaggle competition “Make Data Count – Finding Data References”. Make Data Count (MDC) maintains an open corpus of data citations and this competition seeks state of the art AI advancements (including large language models (LLMs)) to identify mentions of data in papers AND contextualize the relationship.

How did we get here?

April, 2022 – 2024

At the inaugural MDC Advisory Group meeting we agreed the most important advancement for data metrics is an openly available corpus of all data citations, inclusive of datasets with different identifiers in addition to DOIs. This kicked off the work for the Wellcome Trust grant, partnership with Chan Zuckerberg Initiative, and release of the first iteration of the Data Citation Corpus in early 2024. MDC work continued to address the quality of the citations, expansion of the disciplinary coverage of the citations beyond the initial focus on life sciences, and the need for metadata to understand critical facets like discipline, institution, and funder information. 

March, 2024

The MDC Advisory Group discussed the dichotomy of a rapidly evolving tech space external to research versus our MDC world of working with legacy tech and advocacy campaigns targeting publishing workflows to expose data citations. Recognizing the success and reach of Kaggle competitions (e.g., Show US The Data), we applied for both the Kaggle hosting and philanthropic award funds. Our goal: Kaggle participants to harness the latest AI innovations, thereby improving data citation capture. 

September, 2024 – 2025

We hosted a pre-Make Data Count Summit hackathon in London with hopes to develop the training and evaluation dataset for the Kaggle competition. As much as we would have loved for this to be a community developed dataset, labeling efforts require far more than a day-long hackathon. And while standard machine learning approaches for identifying data mentions exist, the unique value of this competition is to programmatically classify the relationship between the data and the paper. Is the data generated as part of the paper (primary mention) or reused or derived from existing records or published data (secondary mention)? This fall and winter we built our training dataset and are proud to release, at the end of the competition, the largest open, labeled dataset of 9k data citations with context. This dataset spans data repositories, disciplines, DOIs and accession numbers.

June, 2025

Today we launch an open and global competition to award 5 competitors who can build an open source LLM that identifies and contextualizes data citations. We are grateful to Kaggle, The Navigation Fund, and Chan Zuckerberg Initiative for their support and are eager to use these models to advance the open data citation corpus for all. Spread the word!

Thank you to Europe PMC for their PDFs and Scott Fisher for his help in preparing the PDF package for this competition as well as Make Data Count team members & advisors Iratxe Puebla, John Chodacki, and Maria Gould for their help in this launch!

Take part in the Kaggle competition