Perspectives on the importance of best practices to incentivize data sharing and data recognition
March 17, 2025 | By: Clare DeanViews from Mike Thelwall, Professor of Data Science
Mike Thelwall, Professor of Data Science at the University of Sheffield, UK, has developed free software and methods for social media sentiment analysis and for systematically gathering and analyzing web and social web data.
Can you tell us a bit about your earlier research on data sharing and data use practices?
One of our earlier projects related to data involved an examination of data sharing in the field of genome-wide association studies (1) – i.e. studies that look for genetic associations with specific diseases. In that field, datasets follow the same format and structure and include good information on data provenance. Data is also often produced by international teams that have to exchange parts within the group and therefore it is already set up for sharing. So we thought this field would be an exemplar in data-sharing practices, and also a good starting point for data usage, as the structure of the datasets makes them easy to combine to generate statistically more powerful results. To our surprise, that did not turn out to be the case. At the beginning of our study in 2010, data sharing stood at 3% for articles from those studies – for this, we recorded whether articles for studies that produced data included information about sharing it; this did increase over time, to 23% in 2017, the last year we analysed. Funder policies requiring data sharing were very likely a driver behind this increase in sharing.
One of my former PhD students, Dr Nushrat Khan, also looked at data sharing and data reuse in biodiversity (2), finding a surprising example of people citing data that they had not used – because it was collectively cited with the data that they had used. Through a survey of data sharing and reuse practices, she also found lots of examples of useful data sharing and reuse (3), and researchers who developed effective support infrastructures to help the process.
The issue of credit often comes up in conversations about data sharing. How do you see practices evolving to enable greater recognition for researchers for the use of their data? What have been the challenges?
Data sharing could be a badge of honour for the researchers. Funders and journals are putting policies in place to encourage data sharing, but if a researcher’s data is used in a later study, there should also be ways for the researcher to claim credit for contributing to that dataset. However, there are a number of challenges until we get to that point. One of them is that a ‘carrot and stick’ for data may not work as a blanket approach across fields. In some fields, the data used for research cannot be easily shared, this can be because the dataset includes human-subject information or GDPR issues, or because the data used comes from a commercial source, making it difficult to share beyond the specific research project.
For bibliometric studies, OpenAlex now provides a database of open information. But even when that research information is open, there is also the issue of the size of the database, how to meaningfully segment the information to the data that is relevant for the study, and whether the program used to extract the data is available.
In the context of journals, many have now implemented data availability statements, however, these are challenging to use as the basis of data sharing and use practices, as the statements are varied and can include a multitude of practices, including that the dataset is not shared. For an evaluation of data usage, looking at data citations is a more meaningful approach.
How do you think the Data Citation Corpus could be a useful resource for research, and in what ways?
I am interested in the Data Citation Corpus and I hope to use it – it’s just that ChatGPT and AI practices have been the focus of my most recent research. Speaking of AI, the intersection between data and AI is likely to bring interesting meta-research topics for exploration in the future. In the area of data, one aspect I’m interested in is disciplinary differences: How do researchers think about what constitutes ‘data’ and does this differ per discipline? What kind of data gets shared and used, and what doesn’t, in different disciplines?
Are there any other tools or information that you think would be useful to facilitate meta-research on data use and data citation?
A couple of key aspects for meta-research in this area are having high-quality data from the papers, and a sufficient sample of data citations – i.e. frequent use of the data citations so these can be studied. It is difficult to achieve currently, but perhaps this is where AI will come in and bring some solutions. I hope AI will be able to identify more data citations from papers, and fill the current gap of citations to data not being part of the reference list or the article metadata. There are also other ways in which data sharing and citation can be improved by AI, for example, publishers could incorporate AI-powered prompts in their workflows to ask the authors ‘is this data in your paper? Will you cite it?’, and facilitate inclusion of more data citations in the articles.
How can we advance recognition for data use as part of research assessment?
It will be important to have examples of success to incentivize data sharing and use practices, especially examples of datasets that have been heavily reused. It would also be useful if we can find examples in as many fields as possible where data sharing and data use led to recognition. As part of the UK Research Excellence Framework (REF), researchers are already invited to note if they have shared their data, it would be useful if they were also prompted to provide evidence that the data was reused, which would demonstrate that their research has had extra value. Showcasing examples of data use evidence would serve as an incentive to others. One of the goals of the REF is to encourage diversity, this also involves diversity in outputs – i.e. you don’t have to only write journal articles, you can write books or software, make a website, or even stage a play. If you do anything research-related well, then you can claim credit for it. That’s the ethos for it and data should be part of it.
References
- Thelwall M, Munafò M, Mas-Bleda A, Stuart E, Makita M, Weigert V, et al. (2020) Is useful research data usually shared? An investigation of genome-wide association study summary statistics. PLoS ONE 15(2): e0229578. https://doi.org/10.1371/journal.pone.0229578
- Khan, N., Thelwall, M. & Kousha, K. Measuring the impact of biodiversity datasets: data reuse, citations and altmetrics. Scientometrics 126, 3621–3639 (2021). https://doi.org/10.1007/s11192-021-03890-6
- Khan, N., Thelwall, M. and Kousha, K. (2023), “Data sharing and reuse practices: disciplinary differences and improvements needed”, Online Information Review, Vol. 47 No. 6, pp. 1036-1064. https://doi.org/10.1108/OIR-08-2021-0423