A recent study published by the British Medical Journal (BMJ) has raised concerns about the reliability of AI-powered search engines and chatbots for delivering accurate and safe drug information to patients. Researchers underscore that a significant portion of the responses pertaining to drug treatments generated by these AI tools were either incorrect or could potentially pose a risk to patient safety.
The study emerged in the context of a technological shift in February of the previous year, where search engines enhanced their functionalities by integrating AI-driven chatbots. This evolution promised users more interactive and comprehensive answers, including those related to healthcare queries. However, despite the expansive data these tools can access, spanning the entirety of the internet, they remain susceptible to producing misinformation or nonsensical and potentially harmful content due to their nature of content generation.
Historically, investigations into the intersection of chatbots and healthcare have focused on the professional perspective, often neglecting the patient viewpoint. Addressing this gap, the researchers embarked on a study evaluating the readability, completeness, and accuracy of AI-generated responses, specifically concerning the top 50 most commonly prescribed drugs in the United States during 2020. This evaluation was carried out using Microsoft's Bing Copilot, a search engine featuring AI chatbot capabilities.
To simulate typical patient inquiries about medications, researchers crafted queries based on common questions patients ask their healthcare providers, such as drug function, usage instructions, potential side effects, and contraindications. A cumulative total of 500 responses were generated through these queries, with 10 questions directed towards each medication under consideration.
The Flesch Reading Ease Score, a metric that estimates the necessary educational level to comprehend a given text, was utilised to assess the readability of the chatbot's responses. Results showed an average score of slightly above 37, indicating that understanding these texts would typically require education at a degree level. Even at their most comprehensible, the texts demanded at least a high secondary school education.
When evaluating the completeness of the information, the chatbot responses averaged a completeness score of 77%, with answers to half of the questions reaching complete accuracy. Nonetheless, responses to a critical question regarding drug precautions averaged just 23% completeness.
In assessing accuracy, expert comparisons with a vetted drug information site revealed discrepancies in 26% of the responses, with 3% totally inconsistent with the reference data. Further scrutiny of 20 randomly selected answers indicated that only 54% matched scientific consensus, while 39% directly contradicted it.
The potential risks associated with following the chatbot's advice were notable: 3% were highly likely to cause harm, and 29% were moderately likely. Alarmingly, 22% of the responses could result in severe harm or even death, though 36% were deemed harmless.
Despite these insights, the researchers acknowledge the limitations of their study, which didn't involve real patient experiences and suggested that results might vary with different languages or in different cultures. They emphasised the importance of seeking professional medical advice as chatbots, despite their technological sophistication, might not always produce error-free responses due to their inability to grasp the full context or intent behind a user’s inquiry.
The research concluded with a cautionary note on the continued use of AI-powered search engines for critical medical information until more reliable and accurate citation mechanisms are developed.
Source: Noah Wire Services