Ten years ago (whilst bootstrapping Portainer) I was a cloud consultant, helping enterprises move to AWS and Azure. I was a true believer, I fully drank the Kool-Aid along with everyone else. Why spend millions on capex, when you could run your IT on-demand, and pay monthly, it sounded awesome!
My team and I completed dozens of migration projects, and all of them had rock solid business cases, promised savings that self-justified the migration, and further promised flexibility and redundancy that an IT exec could only dream of. What was really interesting though is the follow up with the clients 9, 12, 18 months after the migration. All of them struggled to achieve the cost savings, and most of them hadnt really capitalized on the flexibility that being “in the cloud” gave them. Further, conversations with the CFO were becoming heated, as they looked at driving opex reductions to improve the business bottom line. As we all now know, you cant just stop paying for the cloud once you are in it (but you can sweat on-prem hardware!!).
Anyway, I was reminded of all of this last night, in a conversation with a colleague about the current wave of enterprise pushback against AI. This looks eerily familiar to the cloud pushback of 2017, and we already know, roughly, what has to happen next.
Cloud had reached the point where every CIO had to have a story, a migration plan, and every enterprise architect suddenly had a mandate to “get us to the cloud”. Everyone followed the same strategy, “lift and shift”, pick up (figuratively, obviously) the on-prem data center estate, drop it into AWS or Azure, and declare victory. There were even tools to help facilitate the V2C (Virtual to Cloud) migration, conveniently provided by the cloud vendors… of course 😉. The problem was, when you are procuring hardware for you on-prem environment, you are buying sufficient capacity for 3-5 years in the future, and for those running Virtualization, likely they were over-committing their physical server 2 or 3 times in VM allocation. Well, when you pick up those oversized servers, and you move the VMs at 100% allocation, what you end up paying for in the cloud is astronomical. Astronomical^2 when you really think about it.
What followed was the actual work of cloud adoption; the right-sizing, the refactoring, the decommissioning of workloads that never should have moved in the first place, the redesign of applications to use cloud-native services rather than pretend the cloud was a rented data center, and the gradual construction of cost governance so that engineers could not spin up whatever they liked whenever they liked without someone noticing. That work is the reason that today, no serious CIO disputes that cloud is a genuinely good deal for the business. Cloud got good because the industry learned how to operate it, and not before.
AI is right now sitting exactly where cloud was in that first ugly phase. Every board has a mandate, every executive has a story, every enterprise has piled in, and the default move is a version of lift and shift; take the biggest, most expensive frontier model, wire it into every use case that looks like it might benefit from an agent, hand out access to every function that asked for it, and declare progress. It felt good to have AI chat bots everywhere, execs claimed “headcount savings into the millions” and even more interesting, made audacious claims that “our customers are happier with all these bots”. The problem is, now the bill is starting to hammer the living daylights out of business. Month over month the bills keep coming, and increasing. In many cases, the AI spend now surpasses the original cost of headcount it replaced. Rightfully so, the CFO is asking pointed questions about which of these use cases are actually earning their keep, and the same executives who championed the move are quietly discovering that the vendor promises about productivity gains are considerably harder to substantiate than expected.
The pile-in was not stupid; if you were the business exec who did it you are probably reading this defensively, and the pile-in was rational. You had to learn, your organization had to experiment, your people had to build intuition about where the technology helps and where it does not, and the only way to do that is by using it, generously, on real work. The mistake was not moving fast; the mistake was moving fast without governance, without cost controls, and without an operating model, and that mistake is the same one the industry made with cloud, and it is fixable in exactly the same way.
If you are battling with AI costs right now, I encourage you to think about these three points:
Do you know, per workload, which of your AI use cases are producing measurable business value, and which are burning tokens for no return you can point to?
Do you have a governance layer that decides which model is appropriate for which task, so that a customer-support classifier is not being routed through the same frontier model as a complex reasoning workflow?
Do you have a mechanism, agreed in advance, for shutting down AI workloads that do not earn their keep, without it becoming a political fight inside the business?
An answer of “I don’t know”, is totally fine, that helps you pinpoint the work to be done; and it is the same work cloud went through 10 years ago.
The playbook is not special, it’s a simple “rinse and repeat”. Right-size the model to the task (a frontier model being asked to summarize a PDF is the AI equivalent of running a t2.micro workload on an m5.24xlarge). Refactor the workflows that are working so they consume less, and decommission the ones that are not working before they build a constituency inside the business who will defend them. Put a governance layer over the top so that model selection, prompt patterns, and cost per workflow are visible to the people who own the P&L, not just to the engineers running the experiments. And accept that the next act of the AI story inside your business is discipline, not volume.
Every enterprise that went through the cloud cycle did the hard work, most of them after a very direct conversation with the CFO about a very specific line item, and the question sitting in front of you now is whether you do it on your terms, before the conversation is forced, or whether you wait until the conversation is forced and you do it under pressure with a smaller mandate. AI is great for business; AI without a plan, without governance, and without cost controls, is a recipe for disaster.
There is a specific kind of pressure sitting on CIOs and CTOs right now, because it is the reason a lot of this work is not happening yet; every board has been told AI is transformative, every peer at every conference is talking about their AI initiative, and turning around and saying “we need to be more disciplined about how we are running this” feels dangerously close to admitting you got the initial call wrong. You did not get it wrong, you did what the market expected, and you learned what you needed to learn. The next call is discipline, and the executives who make it now, on their own initiative, will be the ones who own the AI story inside their business in eighteen months.
Neil
