Hospitals are embracing artificial intelligence faster than they can check whether it works safely. More than 90% of health systems have deployed third-party AI tools — but fewer than half have the infrastructure to properly test and validate them before clinical use.
A striking 63% describe their AI strategy as still developing or ad hoc, according to industry findings. The risks are specific: model “drift” (an algorithm that grows less accurate after deployment), bias that persists when tools are validated only on generic datasets, and a general lack of continuous monitoring once tools go live. Time, money and talent shortages make structured testing hard to build.
One approach
UPMC offers a template. It built a real-world data platform (Ahavi) to validate algorithms against de-identified patient data before deployment, has run formal AI governance for over two years, tests vendor tools against its own patient population rather than trusting vendor data, and monitors tools after launch. “Without…a dedicated or consistent test environment strategy, they’re going to go down that path just to learn that all that work potentially wasn’t justified,” said UPMC’s Ken Howard.
Why it matters
There are still no consistent industry standards for assessing clinical AI, so health systems and vendors largely self-govern — and clinicians are often left to flag equity or accuracy problems manually. As AI spreads into diagnosis and workflow, the gap between adoption and oversight is becoming one of healthcare’s central safety and business challenges.