A recent study published in the journal BMJ Quality & Safety raises concerns about the reliability of AI-powered chatbots and search engines when it comes to providing safe and accurate drug information. Researchers found that a notable portion of the AI-generated answers could be incorrect or potentially harmful, highlighting the complexity and potential risks of utilizing artificial intelligence in healthcare contexts.

The study, conducted in the context of the considerable transformation of search engines through AI-enhanced chatbots in February 2023, focused on how these technologies respond to queries related to medication. Researchers employed Bing's AI-driven copilot features to analyse responses to the top 50 most commonly prescribed drugs in the United States in 2020. They simulated typical patient inquiries regarding drug usage, mechanisms, instructions, side effects, and contraindications.

The study involved clinical pharmacists and medical professionals specializing in pharmacology to identify frequently asked questions by patients. Each of the 50 drugs was subject to 10 distinct questions, resulting in 500 unique responses generated by the chatbot.

The readability of these responses was assessed using the Flesch Reading Ease Score. The analysis revealed an overall average score of just over 37, indicating that readers would generally need a university-level education to fully comprehend the information. Even the most readable responses required at least a high school education, presenting a challenge for the average patient seeking to understand their medication.

In terms of completeness, while the average completeness score was a commendable 77%, discrepancies were present. Specifically, for five out of the ten evaluated questions, complete responses were achieved, yet the response to a question about precautions when taking the drugs was only 23% complete on average.

Accuracy was another critical aspect evaluated. Of the chatbot's 484 delivered answers, 126 were found to have discrepancies with trusted reference data, while 16 were entirely inconsistent. Furthermore, a subset analysis of 20 particularly low-completeness and low-accuracy answers revealed that only 54% aligned with scientific consensus. Startlingly, 39% of these answers demonstrated contradictions with established scientific understanding.

The potential for the chatbot's misinformation to cause harm was also assessed using the Agency for Healthcare Research and Quality's harm scale. Experts deduced that there was a high likelihood of severe harm in 3% of instances and a moderate likelihood in 29% of cases. Alarmingly, 42% of the assessed chatbot responses were associated with a risk of moderate to mild harm, with 22% bearing the risk of severe harm or death.

Although the study did not use real patient data and acknowledged that language or regional variations might influence chatbot responses, it provided critical insights into the limitations of AI in dispensing reliable medical advice. The study authors noted that chatbots struggled to grasp the intent behind patient queries, leading to errors in the generated information.

The researchers stated, "In this cross-sectional study, we observed that search engines with an AI-powered chatbot produced overall complete and accurate answers to patient questions." However, they raised a caveat, "Chatbot answers were largely difficult to read and answers repeatedly lacked information or showed inaccuracies, possibly threatening patient and medication safety."

The study underscores the necessity of professional medical consultation, emphasizing that AI-powered chatbots should not replace healthcare professionals due to their potential to disseminate incorrect information. As AI continues to evolve, the development of more precise and reliable citation engines remains an essential consideration.

Source: Noah Wire Services