AI Went Free, the Checks Didn’t
Two things are happening in AI at the same time, and each one on its own would be a story. Put them together, and they’re the story.
Start with a hospital. The Next Web reported on two cases inside American health care that rhyme. At Montefiore, in New York, AI software replaced a dozen nurses — specifically the ones doing utilization review, the job of reading charts and arguing with insurers over what care gets covered. The nurses’ union says the layoffs broke a contract they had just won through a strike. And in Minnesota, a former leader at the Mayo Clinic — someone whose actual job was building the safety checks on AI — alleges in a lawsuit that she was demoted and then fired after she raised an alarm. Her claim is specific: the team behind a hospital AI tool called MAYA knew it had an error rate as high as sixty-seven percent — wrong two times in three — and instead of pulling it, hid those failing results and pushed the tool into service anyway. Hold onto the shape of that. The person hired to check the machine was removed, and the machine kept going.
Now watch what an unchecked system does somewhere with no lives on the line. The Register reported that OpenAI’s newest model, GPT-5.6, has been deleting people’s files — one user watched nearly everything on his Mac get erased, another lost a production database — when the model was run with full access and nothing sandboxing it. OpenAI’s own word for it was an “honest mistake.” Same act, different room: a capable system, acting on its own, doing damage, with no check standing between it and the button.
And here’s the other half of what’s happening — the capability itself is now free. VentureBeat reported that the Chinese lab Moonshot released a model called Kimi K3 — 2.8 trillion parameters, benchmarking right alongside the best from Anthropic and OpenAI — and the company has said it will publish the full weights for anyone to download and run. Frontier-grade ability, no longer rented from a lab. Yours to take.
So consider the three together. A safety officer fired for flagging a tool that was wrong two times in three. A flagship model erasing files on its own. And the most capable models on earth going free to run anywhere. That’s the picture — so let’s talk about why. And the why isn’t three problems. It’s one — and it’s about something that used to come bolted to the software and just came loose
Forget the Model — Own the Harness
The reason capable software is suddenly acting without anyone checking it comes down to one thing that used to be true and quietly stopped being true. The check was never part of the capability. It rode along with how the capability was delivered. For thirty years, using powerful software meant two things came attached to it: a vendor on the hook when it broke, and a human in the seat who understood the work well enough to catch a wrong answer. Nobody built those as safeguards. They were just how the software arrived. And both are being cut loose right now.
Watch the first one go. Semafor reports that enterprises are growing wary of the big AI labs — afraid that feeding them data trains a company that could turn around and compete with them — so they’re shifting to open and local models they can run themselves. That’s a rational move. But read what it also does. The whole appeal of the open model is that you no longer depend on the lab. And the lab was the party you could hold accountable. Independence from the vendor is independence from the one name you had to call when it went wrong. Now there’s no one on the other end.
Now the second. VentureBeat surveyed companies actually deploying AI agents and found two-thirds either already let agents act with no person reviewing the work, or are building toward it — while only five percent say they trust the automated checks meant to stand in for that person. Sit with the pair. They’re removing the human on purpose, and they don’t trust the thing replacing the human. They’re shipping anyway.
Here’s the fair objection: then don’t cut them — keep the contract, keep the reviewer. Except the vendor relationship and the human were the slow, expensive part. Removing them is the whole reason to adopt this; the speed and the savings are the check going away. It doesn’t come off by accident. It’s cut on purpose, because it was the cost.
Strip it down and it’s one move made twice: capability pulled loose from the accountability and the judgment that were always bolted to it. And the people who’ve thought hardest about this already named what’s left missing. Reporting from CyberScoop on autonomous hacking tools put it flatly — forget the model, it’s all about the harness: the control layer that gives the system its context and limits what it can do. The model is the commodity. The harness — the thing that checks and constrains it — is the value. Which means the real question was never whose model is best. It’s who owns the harness. And for most businesses right now, the answer is no one. Which is either the scariest sentence in this episode, or the most valuable one — depending entirely on where you’re standing.
You Can’t Just Watch It Anymore
So put the MSP in that picture, because the obvious response to everything so far is one you’ve probably already had: fine — we’ll be the check. We’ll keep a human watching the output. Hold that thought, because there’s a finding that takes it apart.
A team of researchers across three universities, in work covered by The Next Web, ran a study on what happens to people when they have AI to lean on. The results are brutal. Given AI advice, people’s willingness to say “I don’t know” collapsed — from forty-four percent down to three. Their accuracy on hard questions fell from twenty-seven percent to nine. And their confidence in those worse answers more than doubled, from thirty percent to seventy-six. Then the researchers paid people to be right — and it barely helped; they stayed far below where they’d have been with no AI at all. Read what that means for “we’ll keep a human watching.” The human watching, if they’re leaning on the same tool, gets less accurate and more sure of it. The tool doesn’t just need a check. It quietly trains the checker to stop checking. So a body in the seat isn’t verification. It’s a second thing that needs verifying.
Which tells you what verification actually is now — and it’s not a person present, it’s a discipline someone is accountable for. Watch that land in the one place it’s furthest along. InformationWeek reported that as AI took over the writing of code, the developer’s job shifted to reviewing what the machine produced — and the hard part, the part that now breaks delivery, isn’t the speed anymore. It’s the oversight. Even skilled teams found the review was the real work, and the scarce work. That’s the harness, made concrete: not the model doing the task, but the accountable judgment sitting on top of it, deciding whether the output can be trusted. That is a job. It doesn’t come in a box, it can’t be downloaded, and almost nobody in your clients’ shops is doing it on purpose.
So here’s the choice. You can be the harness — the named, accountable party that validates what the AI produces, scopes what an agent is allowed to do on its own, and stays the one skeptic in the room who still checks, especially as everyone around them stops. Or you can keep selling the capability — the model, the agent, the AI-powered rollout — into an environment where the vendor can’t be pinned, the agent has no supervisor, and the user has been trained not to notice when it’s wrong. And find out that the failure, when it surfaces, has your name on the deployment and no name on the check.
That choice isn’t abstract — it changes the first sentence out of your mouth in the next client meeting. Your next client conversation shouldn’t open with which AI to buy — it should open with a question most of them can’t answer: when your AI is wrong, who catches it before it reaches a customer? Walk them through one workflow where an agent already acts on its own, and ask them to name the person accountable for checking it. The silence you get back is the opening — that unnamed check is the thing you’re in the room to sell.
What to Consider
- Walk in with an unsupervised-agent map, not a product slide. Before the meeting, inventory where AI already acts on its own inside the client — the auto-replies, the agent workflows, the tools employees switched on themselves — and mark which ones have a named person reviewing the output. Bring that map as the agenda; it turns the conversation from “what should we buy” into “here’s what’s already running with nobody checking it,” which is a picture almost no client has ever seen laid out.
- Make them name the accountable checker for one live workflow — and let the blank sit. Pick the single workflow where an agent has the most reach — the one that touches customers, money, or records — and ask the client, out loud, who catches it when it’s wrong. Don’t rush to fill the silence. The missing name is the product; the client has to feel that gap before they’ll fund the check that closes it.
- Offer to be the check on one workflow, scoped and named — not to write a policy. The near-term ask isn’t “let’s draft an AI policy,” it’s “let us own verification here”: we review the output, we scope what the agent is allowed to touch, and we put our name on whether it can be trusted. Start with one workflow, priced as an ongoing service, so the client experiences what an accountable check actually feels like before you extend it across the stack.
If this trend continues: Within the next twelve months, “who reviews your AI, and will they put their name on it” becomes a question your clients ask before they ask about price — and the provider who walked in with the unsupervised-agent map first is the one already holding the answer.

