How Do You Tell Real Technology From a Good Story?
Every technology arrives twice. First as a story, and much later as a thing that either works in your building or does not. Telling the two apart is a leadership skill, not a technical one. The question is never whether a technology is real. It is what the technology has been tested to do, who ran the test, and what they had at stake in the answer. The AI Alignment Filter™ runs that question in four steps: name the bottleneck, identify which kind of cognitive load the tool addresses, count the true cost of integration, and decide whether the tool replaces judgment or enhances it. Before any of that, establish one thing. Are you looking at a regulated technology that a disinterested third party has examined, or at software that someone is selling you? Both can be excellent. They demand entirely different questions.
Introduction
I have watched technology arrive in this profession for five decades. It arrives the same way every time.
First there is a demonstration. Someone competent stands in front of a room and shows you the thing working. There is a number on a slide. There is a colleague who already has it and is pleased. The story is coherent, and the people telling it are not lying.
Then, eighteen months later, there is a piece of equipment in a closet that nobody was trained on, a subscription nobody can cancel because nobody is sure what it does, and a team that has quietly built a workaround.
The gap between those two moments is where most technology money goes to die. It is not a gap in the technology. It is a gap in the questions that were asked before the purchase.
This month is about closing that gap, and the first thing to say is that the skill involved is not technical. You do not need to understand how a neural network works to make a good decision about one, any more than you need to understand metallurgy to buy a handpiece. What you need is a way of asking what has actually been established, and by whom.
Key Takeaways
• The useful question is not whether a technology works. It is what it has been tested to do, who ran the test, and what they had riding on the result.
• Regulated technology and unregulated software are not good and bad. They are two categories where completely different amounts of verification work fall to you.
• The AI Alignment Filter™ runs four steps in order: bottleneck, cognitive load, integration cost, and replace or enhance.
• Most adoptions fail at step one, because no constraint was ever named, or at step three, because only the subscription price was counted.
• A tool that replaces judgment fails the Filter no matter how well it scored on everything before it.
• If it does not replace or compress something, it may be unnecessary.
Every Technology Arrives Twice
The first arrival is the story. The second arrival is the Tuesday morning eight months later when someone has to actually use it while the schedule is full.
Stories are not the enemy here. A story is how anyone communicates a possibility, and the people building genuinely useful things have to tell stories too. The problem is that the story and the product are produced by the same people, and the story is finished long before the product is.
So the leader's task is not skepticism. Reflexive skepticism is just as lazy as reflexive enthusiasm, and it costs more over a career, because the person who says no to everything eventually gets passed by things that worked.
The task is to know which questions the story has already answered and which ones it has quietly left open.
The Question Is Not Whether It Is Real
Almost every technology put in front of you is real in the sense that it exists and does something. That is a very low bar, and it is the bar most evaluation stops at.
The questions that separate a purchase you will be glad about from one you will be explaining are narrower and less exciting.
What specifically was this tested to do? Not what can it do, not what will it do next year. What was the tested claim?
Who ran the test? Someone independent, or the company whose revenue depends on the answer?
What were the conditions? A curated dataset in a research setting, or a normal week with interruptions, bad inputs, and a team who did not sleep well?
And what happens when it is wrong? Because it will be wrong, and the honest products have a story about that, while the ones selling a story do not.
Notice that none of those questions require you to evaluate the technology. They require you to evaluate the evidence about the technology, which is a thing any competent operator can do.
The Line That Organizes Everything: Checked by Whom?
Here is the distinction that does the most work, and it is the one to carry into the rest of this month.
Some technology is regulated. Before it can be sold for a stated purpose, a third party with no financial interest in the outcome examines evidence and writes down, in public, what the thing is permitted to claim. You can go and read it.
Most technology is not. It is software, sold on the strength of what its maker says it does, and the only party who has checked is the party being paid.
This is not a moral distinction, and it is not a quality distinction. A great deal of unregulated software is excellent, and regulatory clearance is a floor rather than a recommendation. Clearance means a claim was supported well enough to be permitted. It does not mean the thing will help your practice, fit your workflow, or be worth what it costs.
What the distinction tells you is how much verification work falls to you.
When something is regulated, a portion of the work has been done by someone who was not paid by the seller, and the record of it is public. When it is not, all of that work is yours, and if you skip it, nobody did it.
That is the whole of it. Two categories, two different amounts of homework, and the expensive mistake is treating the second category like the first because the marketing materials for both look identical.
The AI Alignment Filter™
The Filter is a gate you run per tool, before any commitment. Four steps, in order, and the order matters.
Step one: What operational bottleneck is this solving?
Name the constraint in one sentence before you look at any tool at all. If the sentence cannot be written, the search is premature, and everything you find will look promising, because you have no basis on which to reject anything.
A failing answer at this step sounds like "it seems useful." That is not a constraint. That is an impression.
Step two: Which cognitive load does it address?
There are four kinds worth separating: thinking, creating, organizing, and capturing. A tool that reduces the burden of capturing information is doing something different from one that reduces the burden of deciding, and conflating them is how practices end up with three systems that all half-solve the same problem.
A failing answer here is a tool that adds a load rather than removing one, because now there is a new system to learn and maintain.
Step three: What does integration actually cost?
Not the subscription. The training, the workflow change, the ongoing maintenance, and above all the transition period where the team runs the old way and the new way at the same time.
Step three is where most adoptions actually fail, and almost nobody models it honestly. The license fee is never the real cost. The real cost is the eight weeks of doing everything twice, and the person who has to own it on top of their existing job.
Step four: Is it replacing judgment, or enhancing it?
This one is a hard gate, not a consideration. A tool that replaces judgment fails regardless of how well it scored on the first three steps.
A co-pilot, not an autopilot. That formulation has been the position here since the second article I ever published, and it has not needed revising in the eighteen months since.
What a Physician Noticed About the Machines
In Deep Medicine, the cardiologist Eric Topol works through what clinical AI actually does rather than what it is said to do, and one distinction runs through the whole book.
Algorithms perform very well on narrow, specific questions. Is there evidence of this one condition in this one image? On that kind of question, they can match or exceed trained physicians.
They perform markedly worse when asked to assess a whole person. Topol's point is that a clinician reviewing a scan is weighing dozens of possibly related considerations at once, and a system built to answer one question well is not doing that and was never designed to.
He makes a second observation that matters more for purchasing than it sounds. Much of the impressive published performance was produced in laboratory conditions rather than in normal practice, and performance under curated conditions is not a promise about Thursday afternoon.
Topol is not a skeptic. He argues the opposite case: that these tools should be adopted faster for exactly the routine work that is currently eating clinicians alive, precisely so that the humans can spend the reclaimed attention on patients. That is the same exchange this framework is built around. If a tool saves your team thirty minutes of drafting, the gain is not the thirty minutes. The gain is sitting longer with a patient.
He also names a problem the industry has not solved: for the most capable systems, how the machine arrived at its answer remains genuinely opaque, including to the people who built it. In a setting where a clinician carries the liability for the outcome, that is not a footnote.
An Exhibit, Eight Years Old
In 2018, the World Economic Forum published a piece by Alex Gray titled "7 Amazing Ways Artificial Intelligence Is Used in Healthcare." It is worth reading now, and not for its content.
It is a list of confident comparisons. Machines outperforming doctors at detecting skin cancer. A robot passing a medical licensing examination. Software predicting which patients would wake from a coma more accurately than clinicians could. None of the figures carry a citation, a sample size, or a setting.
Read it in 2026, and the striking thing is not that it was wrong. Several of those capabilities are real, and some of them are now cleared products. The striking thing is the distance between the register of the writing and the pace of what followed. Eight years on, the category that has actually been examined and cleared at scale is a narrow one, and it is narrower than the 2018 list implies by an enormous margin.
I am not holding up an old article to score a point. I am holding it up because a document like that one will be in front of you this year, about something else, written in the same register. The tell is not that the claims are outlandish. The tell is that every number is confident and none of them says who counted, or how many, or where.
A Position Worth Being Able to Check
There is a reason I am comfortable saying all of this plainly, and it is not that I have a better forecast than anyone else.
It is that the position has not moved. The second piece I ever published said that AI is not here to replace your team; it is here to make them more effective. When the weekly technology slot began, the caution was already in it: a co-pilot, not an autopilot. The explicit list of things AI must not touch came a year later. The Filter itself formalized it into a gate.
That is an eighteen-month record that anyone can go and read, and I would rather be judged on it than on a framework that appeared fully formed with nothing behind it. In a market where most claims are unverifiable, a checkable one is worth something.
Frequently Asked Questions
Is regulatory clearance a recommendation?
No. Clearance means a specific claim was supported well enough to be permitted for sale. It says nothing about whether the tool suits your workflow, your scale, or your budget, and it is not a statement that the tool is better than the alternative. Treat it as a floor, not a verdict.
Does unregulated mean unsafe?
No. Most business software is unregulated, and most of it is fine. It means the verification has not been done by a disinterested party, so it falls to you. The risk is not the category. The risk is treating the category as though somebody else already checked.
What if the vendor will not answer the evidence questions?
That is an answer. A company with real results is usually delighted to tell you the sample size and the setting, because those numbers are the product of expensive work they are proud of. Evasiveness about methodology is rarely about protecting a trade secret.
How long should the Filter take?
For most candidates, minutes. Step one ends a surprising share of conversations, because the constraint cannot be written down. That is the Filter working, not failing.
We already bought something that fails the Filter. Now what?
Run the same four questions on it anyway and write the answers down. Sometimes the finding is that it is worth keeping and was simply never implemented properly. Sometimes the finding is that it should be retired, and the useful thing is being able to say why out loud rather than paying for it quietly for another three years.
Final Thoughts
The reason this matters is not that technology is dangerous. It is that attention is finite and every system you carry takes some.
A practice, or a company, can only hold so many tools before the tools become the job. That is why the hardest criterion in this whole framework is not impact or cost. It is replacement. What did this retire? If the honest answer is nothing, then you did not adopt a tool; you added a tenant.
Eliminate first. Automate second. Then guard the space you reclaimed, because if you do not decide in advance what the saved time is for, it is absorbed within a fortnight and nobody can tell you where it went.
AI literacy in 2026 is not knowing every platform. It is knowing which ones not to adopt.
Complexity is not innovation.
Sources & Further Reading
Topol, Eric. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again. Basic Books, 2019.
Gray, Alex. "7 Amazing Ways Artificial Intelligence Is Used in Healthcare." World Economic Forum, 2018. Cited here as an example of a register of technology writing, not as a source of findings.
McKeown, Greg. Essentialism: The Disciplined Pursuit of Less. The eliminate-before-you-automate principle in this framework is essentialism applied to tooling, and the debt is his.