The Excellence Problem
We aren't ready for the suspicion required by our new technology.
A family member recently experienced a medical mistake caused by Artificial Intelligence. More than a machine error, the problem stemmed from the conceptual landscape in which the tech was developed.
They saw a doctor of several decades’ interaction. Over time they’ve shared wide-ranging personal histories. While catching up, the physician mentioned an affliction that he, the doctor, had been dealing with for several months. To protect humans privacy, let’s call it a problem of the vagus nerve, which can affect both physical things like digestion and problems of the mind like anxiety. The conversation was recorded and then digitally transcribed by AI.
The patient’s after-visit summary, which arrived a few days later, stated that my relation had a severe and chronic medical condition. After an initial shock, they realized that the AI had ascribed the vagus-type issue to the patient, not the doctor. The physician had reviewed the summary but failed to catch this crucial detail. It would have gone into the patient’s personal health history if the patient hadn’t caught it. Perhaps no big deal in itself, except that in our healthcare system it could have been a tragedy. “Preexisting condition” can affect treatment decisions, or mean bankruptcy when insurers fail to pay.
Thinking about this encounter between two people and a robot eavesdropper raises lots of questions.
Primarily, what was the nature of the error? In Michael Crichton’s novel The Andromeda Strain, a scientist misses key data about a deadly extraterrestrial organism when a flashing red light causes an epileptic trance. In Aldus Huxley’s Brave New World, Lenina Crowne becomes distracted by erotic thoughts and neglects to inoculate an embryo against sleeping sickness. This leads to a tropical death 22 years later. In those fictional cases, the technology is perfect, and the humans are the problem. It’s odd to think that Huxley had more optimistic underlying assumptions about the reliability of tech than we do, but that is the case.
Isn’t the real world more complicated? You bet. In 2009, Air France 447 stalled and plunged into the Atlantic. Everyone died, from a confluence of bad data and information processing, combined with pilot error. Machines failed, while people lost situational awareness and communicated badly. At the Three Mile Island nuclear meltdown, poor information reached the control room and operators acted for hours on a false valve reading, even though there were other indicators that something was wrong. In 1988 the USS Vincennes shot down Iran Air 655 when the crew misread radar tracking data, something that might not have happened under less stressful conditions.
So is the problem that we tend to assume reliability, even excellence, from machines? To some extent. This is an “insufficient suspicion” argument, possibly an issue for my relative’s doctor. At the same time, automation tends to diminish skills and awareness, because we offload attention when we offload work.
Should the doctor be expected to commit additional anxious labor in vigilance over a supposed labor-saving device? That seems not just self-defeating, but impossible. We can scale up our physical and mental capabilities a thousandfold with automation, but inattention is an inevitable byproduct. And such an expectation would be deeply ironic, since the machine is supposed to relieve overwork and stress.
There are many other parts of life where tech has become widespread too fast to be well embedded, and accidents usually become the acceptable level of deviance. Think of the bad facial recognition AI which led to multiple arrests. The cops trusted the AI over their own procedures. They weren’t paying attention (and remember, the false arrests you hear about are only the errors that get caught.) There are lots of stories of people following driving directions to their ruin, rather than trusting what was literally in front of their eyes. Car crashes that begin when a driver is using cruise control happen at higher speeds and result in more fatalities than regular crashes.
So why is there so little push to get rid of these products? Instead, we discuss how they can be made better and, after enough interactions, achieve the goal of flawlessness. The technology is always wonder-working, fast-improving, and perpetually leading to a better overall quality of life. Needless to say, this plays well in the profit-seeking environment where the world’s leading technologies are created, and where ideas that properly belong in sales and marketing often generalize into overall sensibilities about technology.
With the AI medical transcription, or for that matter with facial recognition and bad driving directions, this becomes the new problem: Is doubting our tools now part of the job? That’s not how a doctor relates to a syringe or a blood pressure cuff, since those are reliable and regulated products that work with exceptionally high similarity in each performance. AI errors, however many there are, are probably unique in each case. To catch them all, a doctor risks burnout on a whole new level, in paranoia about his own machines.
On the one hand, today’s tech designers need us to be paranoid; it’s part of their design. That is how products, particularly those with learning loops and over the air software upgrades, are improved. As I wrote elsewhere, AI is usually deployed like most other software has been over the past couple of decades, imperfect at launch and made better in the field. In this case, doctors and patients likely correct the mistakes, and the system updates. But we don’t really know, since tech companies are bad at sharing these details.
Two years ago a medical transcription company called Nabla used Open AI’s Whisper product, which reportedly had error rates as high as eight in 10 uses, and hallucination rates of about 1.4%. Nabla destroyed the original recordings, making audit checks and accountability a daunting task. This was ostensibly for “data safety reasons,” a consumer protection that also helps the company avoid liability issues. In fairness, a company called Abridge, which developed its language model in conjunction with Nvidia, was sued in April for processing information in unregulated sites without patient permission. There seems to be no good system for engineering both privacy and accountability.
Which takes us to the final question: What’s the environment in which this technology is created? Every technology has a larger economic and social context. At one extreme, you might have a situation like science under Stalin, where a researcher might go to the gulag for believing in the evidence of genetic theory. At the other end of the ideology spectrum, a market utterly free of oversight, in which drugs are launched and used with little testing and product safety, sorts out winners (survivors) from losers. Trust in any technology depends to a large extent on trust in the surrounding society.

For more than four decades we’ve had systems that increasingly favor the development of data at speed and with less regulation, which seems to be the way AI is entering even a doctor’s office. The patients given fabricated histories, or the doctor paranoid about his own tools, are like early users on social media who lost privacy in one or another Facebook experiments, or people who searched Google and got weird-looking results. The errors they encountered were steps in corporate learning.
We’re now part of the process of certification, the key elements of learning. Arguably this is a useful process, as long as the stakes aren’t too high, since it’s probably faster than going through a long procedure of standards-setting and regulation. On the other hand, “the hallucinating AI doctor will see you now” doesn’t sound like anyone’s Brave New World.
For this process to work well, suppliers have to act in transparency and good faith, with an open willingness to say that their product is at best provisional, that one should expect mistakes, and that learning from failure will be open and disclosed. Instead, the products are more likely to be released from the market-based contexts in which they’re developed: User restrictions, trade-secret protection, mandatory arbitration, weak liability, deletion of source material, a failure to aggregate reported errors.
In short, development without discipline, and information control that prevents any of the supposedly high-value market transparency. Excellence is promised, but the execution is corrupted by self-interest.



What we have here is a potential AI version of the Satanic Verses.
I discovered an amazing author, Lewis Mumford a couple of days ago. Reading everything he wrote avidly... many conceptual questions that you are asking have answers there.