MADISON, Wisconsin – Researchers at the University of Wisconsin–Madison have raised concerns about the increasing use of artificial intelligence (AI) tools in the fields of genetics and medicine. These tools, while popular, have the potential to lead to erroneous conclusions about the relationship between genes and physical traits, including disease risk factors such as those for diabetes.
Genome-wide association studies (GWAS) are pivotal in identifying connections between genetic variations and physical characteristics. These studies utilise vast datasets, such as the National Institutes of Health’s All of Us project and the UK Biobank, to explore genetic influences on health conditions. However, these databases frequently lack comprehensive data on certain health conditions, which complicates efforts to draw statistically significant conclusions.
Qiongshi Lu, an associate professor in the Department of Biostatistics and Medical Informatics at UW–Madison, highlights the inherent challenges in these studies. “Some characteristics are either very expensive or labor-intensive to measure, so you simply don’t have enough samples to make meaningful statistical conclusions about their association with genetics,” Lu explains.
To overcome the limitations of missing data, researchers have increasingly turned to AI and machine learning models. These sophisticated tools promise to predict complex traits and health risks, even with limited datasets. However, as Lu and his colleagues have discovered, this reliance on AI can introduce biases. They detailed their concerns in a recent paper in the journal Nature Genetics, where they illustrated how a commonly used machine learning algorithm could erroneously link genetic variations with the risk of Type 2 diabetes.
Lu warns about the risk of accepting machine learning predictions at face value. “The problem is if you trust the machine learning-predicted diabetes risk as the actual risk, you would think all those genetic variations are correlated with actual diabetes even though they aren’t,” he notes. Such false positives are a pervasive challenge in AI-assisted studies, transcending specific conditions like diabetes.
To address this issue, Lu and his team have proposed a statistical method aimed at enhancing the reliability of AI-assisted genome-wide association studies. This method is designed to reduce biases introduced by machine learning when interpreting incomplete data, thereby ensuring more accurate predictions. The team successfully applied this method to identify genetic associations with bone mineral density.
In addition to the challenges posed by AI, Lu and his colleagues have identified issues with studies that rely heavily on proxy information. In another study published in Nature Genetics, they critiqued the use of proxy data in bridging information gaps for diseases like Alzheimer’s. For example, some researchers use family health history surveys to infer Alzheimer’s risk, which can lead to misleading correlations between genetics and cognitive abilities.
“Genomic scientists routinely work with biobank datasets that have hundreds of thousands of individuals,” Lu comments. However, he cautions that as statistical power increases, so does the potential for biases and errors in these expansive datasets. The UW–Madison team’s work underscores the necessity for statistical rigor in such large-scale research.
As the integration of AI into genetics and medicine continues to evolve, the research led by Lu highlights the complexities and potential pitfalls of relying too heavily on these technologies without stringent checks. The team's findings emphasize the need for nuanced approaches to ensure the accuracy of genetic link studies, crucial for advancing our understanding of genetics and disease.
Source: Noah Wire Services