
A new study has raised serious questions about artificial intelligence tools designed to predict a person’s risk of stroke or diabetes, after researchers found that some widely used health datasets had no verifiable origin.
The investigation, published in BMC Medicine, examined two popular datasets hosted on Kaggle, an online platform where users share data and machine-learning resources. The files had been downloaded and reused extensively in research. Yet they contained almost no information about who collected the data, where it came from, how it was gathered, or whether it reflected real patients which are important for data analysis.
That missing history matters.
Health prediction models are only as dependable as the data used to build them. If researchers cannot establish where data came from, whether it is accurate, or whether it represents the people a tool is meant to assess, a model may produce unreliable risk estimates. In a clinical setting, those estimates could influence decisions about screening, prevention, referrals or further testing.
The researchers, from Queensland University of Technology and the Australian Centre for Health Services Innovation, found that the two datasets had been used in 125 peer-reviewed studies. Models trained with the data had also been cited in 86 review articles. The research identified evidence that three prediction models based on the datasets had been used in clinical practice. One was cited in a medical-device patent.
The findings do not establish that every study using the datasets reached incorrect conclusions. They also do not show that every model was widely used in routine patient care. However, the lack of basic information about the data source makes it impossible to judge whether these tools were developed using credible, representative information.
That is the central concern.
The study assessed the datasets using the TRIPOD+AI framework, an internationally recognised set of reporting standards for studies that develop or evaluate clinical prediction models using artificial intelligence. It includes expectations about transparency, including details of the source and handling of data. On nine essential criteria relating to data provenance, the datasets scored zero.
Data provenance may sound technical, though the idea is straightforward. It is the documented history of a dataset. It should show who collected the information, when and where it was gathered, how people were selected, what variables were measured, and whether the records were cleaned, altered or combined with other sources. In health research, these details help others assess whether a dataset is credible and whether results might apply to a particular patient population.
Without that information, researchers cannot know whether a dataset captures real clinical patterns or merely appears to do so.
The authors also reported unusual features in the two datasets that raised questions about their authenticity and suitability for clinical prediction research. Their concerns go beyond incomplete paperwork.
Missing provenance can conceal major weaknesses, such as inaccurate records, invented entries, unrepresentative samples or data collected in ways that do not meet ethical and scientific standards.
A prediction model is not the same as a diagnosis. It is a tool that estimates the likelihood of an outcome using information such as age, weight, blood pressure, medical history or test results. A stroke-risk model, for example, may estimate the chance of a future stroke. A diabetes-risk model may identify people who could benefit from further assessment or preventive support.
Such tools can be useful when they are developed and tested carefully. They need to be evaluated in groups of patients beyond the people whose information was used to create the model. Researchers must also show that the tool performs reliably in the setting where it is intended to be used. A model that works well in one hospital, country or population may not work well elsewhere.
When the underlying data is unknown, these essential checks become far more difficult.
The study’s authors warned that models based on data of uncertain provenance should not guide clinical decision-making. Their argument is not that artificial intelligence has no role in healthcare. Rather, it is that healthcare AI requires the same, if not greater, scrutiny as other medical tools. A sophisticated algorithm cannot compensate for questionable data.
An AI system may generate a neat-looking risk score. It may perform well in a technical assessment using the same questionable dataset from which it was built. That does not demonstrate that the score is reliable for real patients. Strong performance figures can be misleading if the original data is inaccurate, incomplete or unlike the people seen in clinical practice.
The research also highlights a weakness in the rapid growth of online health-data sharing. Platforms such as Kaggle can be valuable places for education, collaboration and reproducible research. They allow students, analysts and scientists to share code, test methods and compare approaches. Open data can improve science when datasets are responsibly documented and used within their proper limits.
The problem is not the existence of public data repositories. It is the use of datasets as though they are medically valid when their origins cannot be checked.
Easy access can encourage fast-turnaround research. A dataset may look plausible because it contains familiar variables, such as glucose level, body mass index, smoking status or age. It may be formatted in a way that makes it simple to import into a machine-learning programme. None of that proves it is appropriate for clinical research.
Researchers, journal editors and peer reviewers need to ask basic questions before accepting work built on publicly available data. Where did the information come from? Were the participants real? How were they recruited? Was ethical approval obtained where required? Do the data represent the people for whom the model is intended? Can another team independently verify the source?
If those questions cannot be answered, the resulting model should not be treated as ready for clinical use.
The authors recommended that the two datasets be removed from Kaggle to prevent further misuse. They also called for stronger requirements from journals, research funders and data repositories. Seven articles using the datasets have already been retracted because they were considered unreliable. The findings have also informed updates to the Collection of Open Science Integrity Guides.
Retractions are an important corrective measure, though they do not automatically remove a study’s influence. Published articles can continue to be cited long after they have been withdrawn. Models and methods may be copied into later studies. Review papers may repeat conclusions without tracing the original dataset back to its source. That is why data problems can spread through a research field quickly.
For journals, the case is a reminder that peer review must assess more than the mathematical performance of a model. A study may use advanced methods and still rest on an unsound foundation. Editors and reviewers should require authors to state exactly where their data came from, how it was accessed, how it was processed, and whether other researchers can verify the records.
Funders also have a role. Research grants increasingly support the development of digital health tools, including AI systems for risk prediction, diagnosis and care planning. Funding bodies can require clear data-management plans and evidence that researchers have the right to use the data in question. They can also support independent validation, rather than rewarding only the rapid creation of new algorithms.
Healthcare organisations need to be equally careful. Before an AI-driven prediction tool is introduced into a service, clinicians and managers should know what data trained it, who the intended users are, what outcomes it predicts, and whether it has been evaluated in similar patients. They should also understand its limits.
Patients should not stop taking prescribed medicines, avoid screening, or change a care plan because of this research. The study concerns two specific datasets and the models built from them. It does not mean that every online risk calculator, clinical assessment or digital health tool is untrustworthy.
Still, it is reasonable for patients to ask questions when an algorithm is used in their care. They can ask what the tool is designed to do, whether it supports rather than replaces professional judgement, and how its advice will be considered alongside their symptoms, history and preferences, especially when health AI tools and apps are widely available now. A risk score is one piece of information. It is not a verdict.
The wider lesson is simple. In medicine, transparent evidence is not an optional extra. It is the basis for trust.
Artificial intelligence may help clinicians identify patterns, assess risk and manage growing amounts of health information. Yet the technology must be built on data that can withstand scrutiny. If the origin of a dataset is unknown, confidence in the model should be limited, no matter how polished the final product appears.
The study offers a timely warning as healthcare AI moves from research papers towards real-world settings. Innovation remains important. So does speed. Neither should come before evidence, transparency and patient safety.
The post Study Raises Alarm Over Unverified Data Powering Widely Used AI Tools for Stroke and Diabetes Risk first appeared on PP Health Malaysia.



