Eleven months after we cut 41 people out of a 190-person operations function, I signed the paperwork to bring 17 of them back. Nine returned as contractors at roughly 1.8 times their old fully-loaded cost. I want to be precise about that, because the round numbers people use in conversations about AI layoffs and workforce planning are part of how this keeps repeating: we did not lose 41 and regain 17. We lost 41, spent $1.4M sending them away, spent nine months proving we needed a chunk of them back, and then paid a premium to rent a subset with the institutional memory gone, because the memory had lived in relationships and escalation instincts, not in the runbooks we had them document during their notice period.
Here is the verdict in one sentence, before anything else: AI-driven cuts fail not because the models underperform, but because the automation removes short tasks while the headcount math removes senior people, and almost nobody reconciles those two facts before the list is signed.
I normally write for the people whose names end up on that list. That is the whole reason this publication exists. But I have sat in the other chair too — I built the capacity model that justified our cut, I presented it, and I was wrong in a way that was entirely predictable and entirely legible in our own data three weeks before the decision. Nobody sold me anything; there is no vendor relationship behind this piece and nothing here is sponsored. What follows is the mechanism, the numbers, and the process I use now.
What most people do
The headcount number arrives before the capability map
In my experience the sequence is almost never "we studied what the tool can absorb, therefore N." It is "we need $4M out of opex, therefore roughly 40 heads, and the AI program gives us a defensible story for where they come from."
That second sentence is not stupidity. It is what the operating calendar rewards. The savings target has a date attached and a board deck behind it. The capability study has no date, no owner, and produces an answer that might be inconvenient. So the number gets set in a finance conversation in Q2 and the capability question becomes a downstream justification exercise rather than an input.
I knew this was happening while it happened. I did not stop it, because the alternative was walking into a meeting and saying "I need six months to tell you a number you already have." That is the honest version of how these decisions get made, and it is why the failure repeats across organizations that share no consultants, no vendors, and no leadership.
The cut optimizes cost per head, because that is the column finance can see
Once the target is in dollars, the fastest path to it runs through the highest-cost rows. In a support and operations function, the highest-cost rows are the people with eight to fifteen years of tenure who handle the exceptions, own the vendor relationships, and are the reason your escalation queue does not become a lawsuit.
We did not write "cut senior people" anywhere. We wrote "cut $4.1M" and let the arithmetic do it for us. Twenty-one percent of headcount, thirty-four percent of the function's tenure-weighted experience. Nobody in the room said that second number out loud, because we never calculated it.
The pilot gets graded on the wrong denominator
Our deflection tool was evaluated on ticket closure. It looked excellent. In the pilot it resolved 61% of inbound tickets without a human touching them, and that 61% became the number in every slide, every steering-committee update, and eventually the capacity model.
Here is what nobody ran until month four of the aftermath: those tickets were 22% of handle minutes. Password resets, order-status lookups, shipping-window questions, refund-status checks — the two-to-four-minute tail that makes up most of your ticket count and almost none of your labor.
We removed 21% of capacity against 22% of minutes. On paper that reconciles. In practice it is one of the worst trades I have ever put my name on, and I will explain why in the next section.
Attrition-by-design as a way to avoid deciding
The adjacent shortcut is running the cut through a voluntary program or a hiring freeze and calling it workforce planning. Voluntary exits are self-selecting by external market value. The people who take the package are the ones who can get hired somewhere else this quarter — which is to say, your strongest operators. We ran a voluntary window first and lost four people I would have paid a retention bonus to keep. That was not bad luck. That is what a voluntary program is designed to produce, and I should have known it before I opened one.
What the evidence suggests
Automation eats minutes, not roles, and it eats the easiest minutes first
This is the mechanism underneath most of the reversals you have read about, and it generalizes well beyond support queues.
A language model is strongest on high-volume, low-variance, well-documented work. That is, by definition, the short work. So automation compresses the front of your distribution and leaves the tail untouched. Your ticket count drops hard. Your minutes drop softly. And the composition of what remains gets significantly harder.
Our numbers, twelve weeks after go-live:
- Ticket volume to humans: down 58%
- Human handle minutes: down 19%
- Average handle time on the remaining queue: 9 minutes to 26 minutes
- Escalation rate: 4% to 11%
- First-contact resolution: down 23 points
Read those together. The work that reached a person was now, on average, nearly three times harder than the work that used to reach a person — and we had removed the people best equipped to do it. The tool did exactly what the vendor said it would. Our model of what would be left over was fiction.
The remaining team absorbs the difficulty, then leaves
Within eight months, 12 of the 149 survivors resigned. Nine were people my directors had flagged as regretted attrition. Our voluntary attrition rate in that function roughly doubled against the prior two years.
I do not think this was primarily about grief for departed colleagues, though that was real. It was arithmetic they could feel in their bodies. Every remaining hour was harder than the hour before it, escalations they used to hand to a senior person now landed on them, and the overtime we authorized in month three was an admission that the model had failed without anyone saying so.
Our pulse survey had one item that moved more than any other: agreement with "I trust leadership to be straight with me about changes." It dropped 31 points in that function and 14 points in adjacent functions that had lost nobody. Adjacent. People who were not touched by the decision watched how it was made and drew conclusions about what your word is worth.
The savings do not survive contact with the reconciliation
We projected $4.1M annualized. Fourteen months later, our own finance team put realized net savings at roughly $700K — about 17% of the projection. The gap was severance and continuation of benefits, recruiter fees on the backfills, the contractor premium on nine of the rehires, authorized overtime across three quarters, and the cost of the customer credits we issued when service levels missed contractual thresholds two months running.
That $700K figure is generous. It does not price the 12 regretted resignations or the deals our enterprise team says they lost on renewal because support quality was cited in two of the three post-mortems. I do not have a clean number for that, so I will not invent one.
This is a pattern, not a local embarrassment
You have likely tracked the public versions. Klarna spent a stretch as the marquee example of AI replacing customer service headcount, and later publicly walked a meaningful part of that back and began rehiring human agents after quality slipped — the company's own leadership said the cost focus had gone too far. IBM has been open that AI absorbed a set of HR administrative roles while the company hired more people into engineering and sales. Manufacturers and financial institutions have quietly rehired specialist engineers they released, often through contract vehicles that cost more per hour than the original salary line. Details and figures in these stories move around, so treat any specific number you read as of its moment rather than as current.
The common thread is not that AI disappointed. It is that role-level cuts were made against task-level automation, and nobody reconciled the units.
Three ways the cut gets made
| How the cut gets made | What it optimizes | What it costs you 9–14 months out |
|---|---|---|
| Cost-band cut (highest-salary rows first) | Fastest path to a dollar target; clean to present | Removes the exception-handling layer precisely when automation makes exceptions the majority of remaining work; drives rehires at a contractor premium |
| Flat percentage across teams | Feels fair; avoids internal political fights | Ignores where automation actually landed; over-cuts teams the tool never touched and under-cuts teams it fully absorbed |
| Task-minute cut (map minutes absorbed, then cut where they left) | Matches capacity reduction to real work reduction | Slow — six to ten weeks of instrumentation you may not have; produces smaller numbers than the board wants; still fails if your task data is bad, and most task data is bad |
The third row is what I use now. Note that it gets a negative too. It is not a clever trick that makes the cut painless. It makes the cut smaller and more accurate, which is a harder sell in the room and a better outcome on the reconciliation.
What I actually do now
Inventory minutes, not roles
Before any capacity conclusion, I want the function's work expressed in task categories with volume and handle time attached, and I want to know which categories the automation actually absorbed in production — not in the pilot, and not in ticket counts.
The question I ask the vendor and my own team is one line: what percentage of human labor minutes does this remove, by task category, at production volume? If the answer comes back in tickets, cases, or requests handled, the answer is not usable. Most vendors will give you this if you insist. The ones who cannot are telling you something.
Run shadow mode through a full cycle, including a peak
We went live on a projection built during a low-volume stretch. Now I want the automation running in parallel with the existing team for at least one full business cycle including a seasonal peak, with humans reviewing the automated output rather than the automation running unattended.
Shadow mode costs money and it delays the savings. Say that out loud in the meeting rather than hiding it. The trade is a capacity number built on observed production behavior instead of a pilot conducted on the friendliest possible sample.
Run the residual-difficulty test before you choose names
Take the work the tool absorbs. Remove it from the queue on paper. Now describe what is left: average complexity, escalation likelihood, tenure typically required to resolve it, how much of it depends on relationships or undocumented context.
If residual difficulty goes up — and it almost always goes up — then a cost-band cut is structurally wrong for that function, and you need to be able to say so with the specific number in front of you. My version of this is a single slide: current mix versus post-automation mix, with tenure-to-resolve on the y-axis. It took an analyst four days to build. It would have changed our decision.
Protect the escalation layer by name
I now identify, before any list exists, the people who are load-bearing for exceptions, regulator-facing work, and vendor relationships. Not "high performers" in the review-rating sense — load-bearing in the operational sense. Who gets called when the thing breaks at 11pm and the runbook does not cover it.
Those names go on a protected list that the cost model is not permitted to touch. If that makes the target unreachable, then the target is unreachable, and that is information the CFO needs in June rather than in April of the following year.
Put re-acquisition cost in the business case as a line item
Every cut proposal I sign now carries an explicit estimate: if we are wrong by 30%, what does it cost to reverse? Severance already spent, recruiting, the contractor premium, ramp time to competence, and the probability that the specific people are gone. For us, reversing 17 roles cost roughly 2.4 times what retaining them would have.
The purpose of that line is not to prevent the cut. It is to price the confidence interval so the room can see what it is betting on.
Redeploy before you remove
When automation absorbs a task category, the person who owned that category has just become the most qualified person in your organization to supervise, audit, and improve the automation of it. That is a real job with real value, and it is cheaper than backfilling the capability from outside in fourteen months.
We converted six roles this way in the second round. Four worked well. Two did not — the people did not want the work and said so, and I would rather know that in week two than pretend redeployment is a universal answer.
The five questions I will not sign without
- What percentage of human labor minutes does this remove, by task category, at production volume?
- What does the remaining work look like after the easy tail is gone, and who is qualified to do it?
- Which specific people are load-bearing for exceptions, and are any of them on this list?
- What does reversing this cost if we are wrong by 30%?
- What will the people who stay conclude about how this decision was made?
That last one is not a soft question. It priced out at 12 regretted resignations and a 31-point trust drop in our case, which is a larger number than most of the hard ones on the list.
The honest cost of doing it this way
This approach is slower by six to ten weeks. It produces smaller cut numbers, which means you will be in a harder conversation with finance and possibly with your CEO. It requires task-level data that many organizations do not have and cannot produce quickly. And it does not always change the answer — sometimes the automation genuinely absorbed the work, the cut was right-sized, and all you bought was confidence.
Confidence is worth the six weeks. But I am not going to tell you it is free, and I am not going to tell you it will save every job. It will not.
Who this is for, and who it isn't
This is for the executive who owns or meaningfully influences both the number and its sequencing — who can go back and say "the capacity model is wrong and here is the specific reason," and be heard. If you have that standing, the residual-difficulty slide is the highest-leverage four days of analyst time available to you.
It is also for the leader currently nine months post-cut, watching service quality slide and wondering whether to admit it. Reversing at month nine costs less than reversing at month eighteen, and the survivors already know. The only person the delay protects is you.
It is not for someone who has been handed a fixed number and a Friday deadline with no discretion. In that position, the one thing still inside your control is which roles absorb the cut. Protect the exception layer, take more from the categories the automation actually emptied, and document your objection in writing so the reconciliation has an author. That is a smaller move than I would like to offer you. It is still worth more than a better-argued objection you never make.
And it is genuinely not needed where automation absorbed a workflow end to end — where a task category dropped to near zero minutes and stayed there through a peak. Those cases exist. They are rarer than the deck implies, and they look completely different in the data: the minutes leave, and they do not come back.
Back to the 17
I opened with 17 rehires because that was the number my board asked me about, and for a while I understood it as the measure of how badly we had failed.
It is not. The 17 were the receipt, not the failure. The failure was a single unreconciled unit: we graded the automation in tickets and cut the humans in dollars, and those two numbers never met in a room before the list was final. Everything after — the overtime, the credits, the resignations, the contractor premium — was the invoice for that one omission arriving in installments.
Seventeen rehires is not the price of getting AI wrong. It is the price of never asking which minutes the machine was actually taking.