Build vs. Buy vs. Operate
(Contributed by Luke Southwell-Chan, VP Strategy)
Early in my career I worked at a company building safety-critical AI systems. The compute demands were enormous even by pre-ChatGPT standards, so the company invested in on-premises infrastructure, hundreds of GPUs, a dedicated facility, a strong team to run it. Over time the operational burden of that environment grew until it began competing with AI development itself for engineering talent and budget. So the company made what felt like the rational move and shifted significant workloads to the cloud. The complexity did not disappear. It changed shape. The problem had not been solved. It had been translated.
That experience is why I believe the entire build-vs.-buy framing is broken. Not because either option is wrong, but because both options leave the same question unanswered: who operates the environment once the decision is made?
The Old Framing
Most enterprises running production AI are working through some version of this analysis right now. They estimate the cost of doing it themselves versus paying a cloud provider, finance runs the model, procurement runs the process, and a decision gets made. The framework works fine for traditional enterprise IT, where the workloads are predictable and the operational surface is well understood. AI infrastructure is neither of those things. Training runs consume massive clusters for days and then release them. Inference requires sustained, low-latency capacity that cannot tolerate contention. Compliance obligations in regulated industries extend into the infrastructure layer itself. The old framework was never designed to evaluate workloads like these, and it shows.
Building gives you control. For certain organizations at certain stages, that matters more than anything else. But the limit of building is not the build. It is the operate. Running a production GPU environment is a permanent, specialized function that draws from the same talent pool as the AI programs it supports, and it never stops growing. Most enterprises that go this route discover within two years that the operational cost, measured in engineering bandwidth and organizational attention, is significantly larger than what the original business case projected. The hardware was the easy part. The ongoing operations are what strain the organization.
Buying should solve this, and for general-purpose compute it largely does. The hyperscalers have built global, deeply integrated platforms that are foundational to how modern enterprises operate, and any serious AI infrastructure strategy needs to work with them, not against them. But "managed" in the cloud context typically means the provider handles the infrastructure layer while the enterprise handles everything above it: scheduling, monitoring, compliance controls, cost governance. For production AI at scale, that residual surface is not small. There is also the question of capacity. Enterprise GPU demand has outpaced available supply, and the gap is structural rather than cyclical. The compute enterprises need either does not exist in the quantities required, cannot be ringfenced to a single tenant, or comes with lead times measured in quarters. Many enterprises buying cloud compute for AI still end up staffing significant internal operations teams, which is the same organizational cost they were trying to avoid by not building.
The Third Question
Both paths assume the enterprise will carry operational burden. The only variable is who owns the hardware and where it sits. The question that changes the framework entirely is what if the enterprise does not operate it at all. Not outsourcing to a systems integrator. A dedicated, single-tenant environment, purpose-built for the workload, integrated with the cloud platform already in use, and operated end to end by a team whose sole function is keeping it running. This is not managed services, which reduce the operational surface. A fully operated model eliminates it altogether.
The distinction matters because the operational burden of AI infrastructure is not a line item you can optimize through better tooling. It is a structural tax on the organization. It consumes engineering headcount, it entangles every AI initiative with an infrastructure dependency, and it creates drag on the strategic roadmap that accumulates quietly over time. When that tax lifts, the engineers who were managing infrastructure go back to building models. The decisions that were bottlenecked on cluster capacity get made on business merit. The roadmap stops eroding.
I think about that early-career experience often, not because anyone made a wrong call, but because the right framing did not exist yet. The choice was build or buy. The actual question, who should operate it, was never on the table.
Luckily for enterprises today, it is on the table now.
More posts
The Missing Fourth Option
Every major computing platform reaches the same moment: enterprises stop buying components and start buying operating environments. AI infrastructure is at that moment now. Three current options don't fit. Here's the missing fourth.
Why we built the Enterprise OS Cloud
Most companies raise to build a category. Dapple built it, validated it, then raised. The founding thesis behind the Enterprise OS Cloud.
Beyond Cost Per GPU: The Metric That Actually Matters
Engineering leaders should not have to choose between looking cost-efficient on a slide and actually delivering for the business. That choice is a symptom of measuring the wrong thing.