Home Top Stories The First Rule Of Healthcare AI
Top Stories

The First Rule Of Healthcare AI

Share
The First Rule Of Healthcare AI
Share

At heart and by training, I’m a computer scientist, not a clinician. But some of my most important lessons about healthcare AI came from healthcare settings—like the VA hospital in North Chicago. This is where my team and I ran our first big test of an AI system built to manage patient data. With a system developed for managing that data, the next question was: How do we test it? In testing the answer, we learned the first rule of healthcare AI: It’s all about clinical data accuracy.

In Pursuit of a Faster APACHE Score

In the early 1980s, I helped lead a project centered on the Acute Physiology and Chronic Health Evaluation (APACHE) score, at the time the most well-known scale for determining the seriousness of an acute illness. It was essentially used as an indicator of mortality rates in ICU patients; the higher a patient’s APACHE score, the higher the risk of mortality. However, it couldn’t be calculated until at least twelve to twenty-four hours after admission.

Our goal became to build a model that could get ahead of that lag, digitizing doctors’ clinical judgment and predicting a patient’s trajectory before the APACHE score was calculable. However, we knew that the model we were building would only ever be as trustworthy as the data clinicians actually recorded. The question was whether we could capture that judgment accurately enough for a machine to learn from it.

Teaching a Model to Think Like a Doctor and Prioritize Clinical Data Accuracy

With the help of physicians at the VA hospital in North Chicago, we identified the different clinical indicators doctors consider when evaluating patients and came up with approximately 1,500 rules or descriptions that doctors use for patient assessment, going beyond the standard APACHE criteria. We then obtained records from ICU patients who had already been discharged or passed away, and had medical students digitize one hundred ICU patient cases with time-stamped clinical events.

Since it would have been unrealistic to expect the clinical team to assign probabilities to each of the 1,500 clinical findings, we selected ten charts out of the one hundred to serve as our training set. We had three doctors review those ten training charts in detail. For every clinical event in each patient’s stay, the doctors scored it on a scale from one to seven, indicating severity level. We then converted those scores into probabilities and used them to train a pattern-recognition AI, feeding the results into a Bayesian model.

The breakthrough came when we discovered that our model could predict the twenty-four-hour APACHE score by the fourth hour, a dramatic improvement that correlated directly with most providers’ assessments. That result was only possible because the doctors’ judgment was carefully captured and translated, allowing for the necessary clinical data accuracy. Without accurate data, the model would break down. This rule still applies to today’s LLMs and clinical-AI builders.

Four Decades Later, the Same Principle Applies

Today’s healthcare AI can hold huge amounts of medical terminology, and its capabilities keep expanding. But as AI evolves and goes beyond regurgitating knowledge to generating it, the risk grows as well. The standard medical terms clinicians rely on may become distorted or polluted.

To circumvent this risk, the same discipline regarding clinical data accuracy is needed that we relied on in that ICU decades ago. Done right, human insight and AI work in harmony—but only if we humans stay in control of the language and the systems that carry clinical meaning between us.

Source link

Share

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *