AI Search Tools Misidentify Validated Blood Pressure Monitors in Preliminary Study
Researchers found errors even when tools located devices on official validation registries. Answers also changed when some questions were repeated.
Loading page…
Researchers found errors even when tools located devices on official validation registries. Answers also changed when some questions were repeated.
Listen to this story
In a preliminary conference abstract, researchers reported that four AI search tools did not reliably distinguish validated home blood pressure monitors from unvalidated ones. Across 324 devices tested in Canada in April and May 2026, accuracy depended on the wording of the question; Gemini scored 86%–91%, while ChatGPT, Copilot and Perplexity results ranged from 63%–83%. The results are a snapshot of tools at that time, not a peer-reviewed study or a measure of current performance. Patients and clinicians should verify devices in recognized registries rather than rely on an AI response.
The sample included 145 validated and 179 unvalidated monitors, drawn from public registries, a list of known unvalidated devices and top sellers on Amazon in three countries.
All four tools were less accurate at recognizing validated monitors, including lesser-known devices listed in official registries.
When researchers repeated questions for 20% of devices with mixed results, the same tool often answered differently across examiners, computers and days.
Asking an AI search tool whether a home blood pressure monitor has passed accuracy testing can produce the wrong answer—even when the device appears on an official registry. A preliminary study of ChatGPT, Microsoft Copilot, Perplexity and Google Gemini found that all four sometimes rejected monitors that had already met clinical validation standards.
The American Heart Association published the findings October 8, ahead of the study’s scheduled presentation that evening at its Hypertension Scientific Sessions 2026 in Arlington, Virginia. The research is a conference abstract, not a peer-reviewed paper, and its findings remain preliminary.
Researchers tested 324 monitors in Canada during April and May 2026: 145 validated devices and 179 unvalidated ones. The sample drew from StrideBP, ValidateBP and Hypertension Canada registries, a list of known unvalidated devices, and top-selling monitors on Amazon in Canada, Australia and the United States.
Each tool received three question formats for every device. The researchers varied how much direction the prompt supplied, rather than testing only one casual question:
Google Gemini performed best. Accuracy varied with question phrasing, and the reported ranges below should not be read as a single score shared by every product.
Gemini’s correct answers ranged from 86% to 91%, depending on question wording.
Results for the other three tools ranged from 63% to 83% across tools and question formats.
All four tools were less accurate at recognizing validated monitors than unvalidated ones. That difficulty extended to lesser-known devices, even though those monitors appeared on the same official registries as more popular products.
Presenting author Anna Soriano, an internal medicine resident at the University of Montreal, said the tools sometimes found an official listing but appeared unable to interpret it as proof of validation. The researchers could not determine exactly why the tools struggled with devices that had passed testing.
Consistency was another problem. Researchers retested 20% of devices that had generated mixed results, using three examiners on different computers and days. The same tool often gave a different answer to the same question. The April–May testing also makes these results a historical snapshot, not a measurement of today’s tools.
Validation means independent, third-party testing has shown that a monitor gives consistently accurate blood pressure readings. The American Heart Association’s 2025 guideline recommends using a validated home monitor. Researchers advised patients and healthcare professionals to confirm a device’s status through recognized, free public registries.
Soriano warned that an incorrect AI answer could lead someone to believe an unvalidated device had passed testing. Inaccurate readings from such a device could then contribute to inappropriate diagnosis or treatment decisions. That was a potential consequence she identified, not a patient outcome measured in this study.
Loading discussion...
Join the conversation
Explain when a direct source should replace an AI answer.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.