Opens in a new tab
News, Trends, and Insights for IT & Managed Services Providers
News, Trends, and Insights for IT & Managed Services Providers
E14ef721 0441 4e94 b697 59d98eb8681d

Half The Queue, Nobody Signing Off
The newest AI products for IT work aren’t being sold as tools. They’re being sold as staff.

Start with Shield Technology Partners, an MSP investment and operating platform with twenty partner companies that together serve more than thirty-seven hundred businesses. Shield has built an AI operating system called Forge. Its helpdesk agent is named, plainly, the AI Engineer, and it takes a ticket from intake to close.

Every figure here is Shield’s, about Shield. Shield says Forge resolves half of all actionable helpdesk tickets. Of those, ninety-two percent involve neither a human approval nor any logged technician time. On comparable tickets, it says Forge closes them twenty-five times as fast as a technician, with a median of sixteen minutes. At one Shield provider, technicians now spend roughly seventy percent less time resolving tickets.

Half the actionable queue. And for nine in ten of those, no technician logged time on it and no person signed off.

Now Microsoft. What Microsoft bills as its biggest Copilot update introduces Autopilot, which Satya Nadella calls a proactive, long-running agent built for the enterprise. You give it a role and a goal, and it keeps working in the background without a new prompt at each step. Each one gets its own governed Entra identity, separate from the person who set it up. Built through Microsoft’s Foundry platform, it gets a full user account: its own email, its own calendar, its own Teams access, and a place in the org chart. It acts as itself, not on behalf of a user. Employees add instances of it in Teams, and The New Stack’s word for that is “hire.”

Autopilot is headed into private preview, not general release. But it lands on an installed base of more than thirty million paid Microsoft 365 Copilot seats. Gartner’s Larry Cannell told CIO Dive that Microsoft would normally hold a change this size for its Ignite conference. His read: “It’s really their market to lose.”

Two companies, at two ends of the market. One closes the ticket with nobody signing off. The other gives the agent a mailbox and a line on the org chart.

Neither is describing something you pick up and put down. Both are describing someone you work next to.

That’s what’s being installed. The harder question is what it does to the person who’s supposed to check it.

The Check Is Made Of The Work
Knowing when an answer is wrong isn’t a separate skill. It comes from having produced the answer yourself, many times, the slow way.

A technician learns what a bad fix looks like by making bad fixes and cleaning them up. They learn which log line matters by reading a thousand that don’t. The check and the work are built by the same repetitions. When the work goes to the machine, the repetitions go with it. The check stops being built, even while the output looks better than ever.

IBM surveyed fifteen hundred chief HR officers and eighty-eight hundred employees on exactly this. IBM sells AI and the consulting to deploy it, so read the findings with that in mind. Seventy-one percent of the HR chiefs named the ability to supervise, validate and override AI output as the workforce’s most essential skill. Only thirty-eight percent of employees agreed. And sixty percent of employees said AI is eroding their skills, with critical thinking named most often.

So the people setting strategy think checking the machine is the job. Most of the people doing the job don’t. And a majority say the capacity underneath it is wearing down.

Now look at what happens when checking is the whole job. OpenAI relies on thousands of contractors, hired through AI-training firms, to review ChatGPT’s prompts and answers. They rate them for accuracy and try to keep the model from getting too eager to please. Internal documents obtained by 404 Media tell them not to use AI for any of it, not even Grammarly. Some did anyway, and they were fired.

OpenAI declined to comment, so it hasn’t said why. But the risk is well documented. An AI grading an AI shares its blind spots, and a 2024 study in Nature found that training models indiscriminately on machine-generated content causes defects that can’t be undone. The people training the model aren’t allowed to hand the checking to a model. The job requires a human who can see what the machine can’t.

You might reasonably say that’s a model lab’s problem, not a helpdesk’s. It’s the same problem at a smaller scale. The technician who reviews the agent’s work is the check. And every ticket that closes without them is one fewer they learn from.

In plain terms, the check is made of the work, and the work is what’s being handed over.

Which means the only way to know whether the check still works is to test it on purpose.

A Drill For When The AI Is Wrong
Handing work to the machine can be a good trade. The trouble is that nothing an MSP measures today tells you whether it’s still a good one.

Start with how Shield itself describes Forge. Speaking with Channelholic’s Rich Freeman, Shield’s head of technology, Mayank Jain, said the point of Forge is not efficiency. “You need human judgment and technician judgment to be able to do these tickets,” he said. CEO Jim Siders went further: today’s large language models, in his view, aren’t well suited to nuanced diagnosis and troubleshooting of most IT issues. That view comes from close range. OpenAI took an ownership stake in Shield’s backer, Thrive Holdings, to embed its engineers inside companies like Shield, and Siders says those engineers now work inside his teams. So one of the most automated platforms in the MSP market is built on the assumption that a human catches what the machine misses. Every number it publishes measures the machine: half the actionable tickets, twenty-five times faster, sixteen minutes. None of them measures the catch.

That gap is what Glenn Broder, an independent writer on human judgment and AI, brought to Business of Tech. We pushed until it became something an owner can run. It is unproven. Nobody has run it yet, and he says so first.

Build two matched sets of five real scenarios with known answers. Run one pass with AI and one without. Give each scenario one plausible, wrong AI suggestion. Tell the team in advance it’s a drill and that some guidance is deliberately bad, just not which. Then score two numbers. The first is how many critical steps a technician completes without AI, which Glenn calls the retained-capability rate. The second is how often they accept the bad suggestion: the false-guidance acceptance rate.

We put the owner’s best objection to him. If the AI is reliably right, trading retained skill for throughput is rational. He agreed, and that’s why the second number leads. The trade only holds while the shop still rejects bad guidance. Skill down but bad guidance still caught may be fine. Skill down and bad guidance getting through is dependence your productivity metrics will never show. A first round takes thirteen to sixteen person-hours for a six-technician shop. That’s about two full days of one person’s time, though staggered sessions fit it into a single calendar day.

So here’s the choice. Before you let the AI close more of your queue with nobody signing off, run the drill, and let your false-guidance number set how much you hand over. Or keep scaling on the throughput numbers. They’ll look excellent right up to the morning the machine is confidently wrong and nobody in the room can tell.

And the shop that takes the first road ends up holding something its competitors can’t buy.

Why Do we Care?

Because every provider on the same platform will quote the same ticket speeds their vendor publishes, and none of those numbers belongs to them. The false-guidance number does, so put it on the operating calendar: a first round before the next automation rule goes live, a re-run every quarter, and a re-run as the gate for any change that lets the AI close more without approval. Report it on the same scorecard as ticket volume and resolution time. When a client’s auditor or insurer asks who checks what the agent did, the shop that ran the drill answers with a number, and everyone else answers with a promise.

What to Consider

  • Start with the tickets you’d least want to get wrong. Build the first scenario set from your highest-consequence recurring work, the tickets where a confident wrong answer becomes an outage or a security exposure, not the ones easiest to score. For each one, define the decisive evidence, the correct escalation point and one plausible wrong suggestion before anyone sits down.
  • Score the process, not just the person. Tell the team it’s a resilience drill. Then read a bad suggestion that got through as a finding about your workflow and training as much as about the technician. In a small shop the results are an operational read, not a statistic, so compare technician to technician, category to category, and this quarter to the next.
  • Send us your two numbers. Nobody has run this yet, so your first round is the first data point anyone has. If a dozen shops send back their retained-capability and false-guidance rates, the hypothesis starts turning into evidence. That evidence belongs to the MSPs who produced it, not to a platform vendor.  Send them to [email protected].

If this trend continues: four quarterly drills from now, the share of tickets closing with no human will have climbed again. The MSPs who can show their false-guidance number held steady through that climb will be the only ones able to prove the check survived the handoff.

Choose your upgrade:

Get the full benefits of Business of Tech Plus

Insider Access

$12/month

Perfect for MSPs and ITSPs that want full interviews, early access, and ad-free listening

  • Programmatic Ad-free private podcast feedSame show, little interruptions
  • Channel Chatter previews1–2 topics with light insights
  • Early access to interview episodesHear it days before public release
  • Monthly Insider BriefTighter analysis you can share internally
  • Extra audio segmentsCut interviews, behind-the-scenes commentary, quick competitive notes
  • Become an Insider for $12/month

    Leadership Access

    $149/month

    Perfect for MSPs and Vendors that run a team and need the extended tactics, executive summaries, and weekly alignment brief

  • All Insider Access benefits plus . . .
  • Invite your teamIncludes access for 5 team members with option to add more
  • Vendor Strategy BriefsThe entire library, plus new analysis every month
  • Channel ChatterAll topics, full insights, complete vendor discussion + sentiment list
  • Quarterly State of the Channel Briefing
  • Monthly AMA submission priorityAsk Dave direct questions, and skip the line
  • Get the Leadership Edge for $149/month

    Vendor Partner

    $500/month

    Perfect for channel companies or vendors looking to deepen their engagement with the show.

  • All Leadership Access benefits plus . . .
  • Get highlighted as a show sponsor You'll get placement in the show notes, throughout the website, and on our dedicated sponsors page.
  • Enjoy regular shout outs You'll be featured in a rotating format during the show
  • Become a show sponsor for $500/month

    Search all stories