Assessing the Reliability of AI Health Chatbots: ChatGPT Health and Beyond
In a bid to revolutionize health information access, OpenAI recently unveiled a health-focused space within ChatGPT. The new feature promises a safer platform for discussing sensitive topics such as medical data, illnesses, and fitness. However, an analysis of ChatGPT Health's effectiveness raises concerns about its ability to deliver accurate and reliable insights.
Inconsistent Health Grades
A report by The Washington Post revealed that ChatGPT Health's assessment of a user's cardiac health could be highly inconsistent. When provided with a decade's worth of Apple Health data, the chatbot assigned an 'F' grade to the reporter's cardiac health. However, a subsequent review by a cardiologist deemed the assessment baseless, stating that the reporter's actual risk of heart disease was extremely low.
Further tests showed that ChatGPT Health's grades for the same data could fluctuate significantly, ranging from an 'F' to a 'B' across conversations. The chatbot occasionally ignored recent blood test reports it had access to and occasionally forgot basic details like the reporter's age and gender.
Overreliance on Smartwatch Metrics
The inconsistency in ChatGPT Health's health assessments can be attributed, in part, to its overreliance on smartwatch metrics. The tool leans heavily on Apple Watch estimates of VO2 max and heart rate variability, which have known limitations and can vary significantly between devices and software builds.
Independent research has found that Apple Watch VO2 max estimates often run low. Yet, ChatGPT still treated these estimates as clear indicators of poor health. This reliance on potentially unreliable data could lead to inaccurate health assessments and misguide users.
Implications for North East India and Beyond
As health consciousness grows in North East India and across India, the appeal of AI health chatbots is likely to increase. However, the inconsistencies and overreliance on smartwatch metrics highlighted in the analysis of ChatGPT Health underscore the need for caution when interpreting the insights provided by these tools.
While AI holds immense potential for unlocking valuable insights from long-term health data, early testing suggests that feeding years of fitness tracking data into these tools currently creates more confusion than clarity. Users are advised to approach these tools with a critical eye and consult with healthcare professionals for accurate and personalized health advice.
Looking Forward
The development of AI health chatbots is still in its infancy, and it is essential for companies to address the issues highlighted in the analysis of ChatGPT Health to build trust and ensure the safety and effectiveness of these tools. As AI continues to evolve, it is hoped that these tools will become more reliable and accurate, providing valuable insights to help users maintain and improve their health.