The Multiplier and the Boundary

(Contributed by Jérémy Tetu, AI Infrastructure Engineer)

At Nokia, I worked on a CLI tool that bootstrapped and managed Kubernetes clusters. Bare metal. OpenStack. AWS. Whatever ran underneath, the same command spun up the cluster.

The turning point wasn't "this is faster." It was that the cluster became a declared state instead of a sequence of steps to replay by hand. Two separate runbooks became one command.

Two ways to get it wrong became one way to get it right.

That's the moment when automation became the architecture itself. The product was no longer a cluster. It was the ability to get any cluster, without relearning the layer beneath every time.

An ability multiplies. A runbook doesn't. That's the definition of a force multiplier: the same human effort covers one cluster or twenty without changing in nature.

Real automation is certainly harder than it looks. And there are four failure modes I've watched teams fall into again and again:

Automating the wrong layer. Scripting the symptom instead of fixing the root cause: a cron that restarts a crashing pod. The failure signal disappears. The failure stays.

The fake-automated. A human running a twelve-step runbook from memory is not automation.

It's a human executing a script in their head. My test: does it work the same at three in the morning without the right person being awake? If the answer is no, it isn't automated. It's documented.

Over-automating. The urge to automate everything leads to a system that costs more to maintain than the manual gesture it replaced. Change now means touching the whole contraption. The multiplier becomes a divider.

Under-automating the boring stuff. Provisioning, certificates, secret rotation, opening network flows. The tedious repetitive work that eats the time that should go toward real engineering. We under-automate the tedious because it isn't rewarding to automate. That's a misallocation.

The discipline that avoids all four is straightforward on paper and hard in practice.

Idempotence, non-negotiable. Observable state matters more than the execution log - that's the difference between a script and a controller. The failure path is designed before the happy path. You test it by rerunning to convergence and killing a node on purpose. The multiplier is decided upstream, not once you're already scrambling underwater.

This matters most at AI infrastructure scale. A training run ties up hundreds of GPUs for days, with no human watching continuously. Topology matters as much as raw capacity.

Inside an NVLink domain a GPU can talk to its peers at roughly 900 GB/s aggregate. Cross that boundary onto InfiniBand and it's in the range of 50 GB/s per GPU. Different bases, but the gap is an order of magnitude however you count.

Thermal is worse because it's silent. In a synchronous data-parallel job, every GPU barriers at each allreduce - they wait for one another. So one GPU that throttles doesn't just slow itself down. It becomes the straggler that paces the entire collective. Hundreds of GPUs idling at the barrier for one that's degraded. Nothing crashes. The run just quietly takes longer.

That's where the difference with traditional infra is sharp. In classic infra, degradation shows up fast: latency, 500s, an alert. In AI batch infra, it shows up as "the run took 30% longer." Invisible without automation that watches topological health and reschedules, not just uptime.

Every PB-scale workflow becomes a distributed systems problem. Every distributed systems problem requires automation to see what a human structurally can't at this scale and duration.

At ITSF, we grew from one customer environment to ten while the team grew from one engineer to four. The first environment took two of us a month to stand up. By the end, one person did it in five days. Small team, same work. Because the automation was doing the repetitive work in our place.

At Dapple, that discipline is the foundation. I design our sandbox platforms - prod-like environments anyone can stand up on command, emulating a datacenter, simulating the GPU fleet, producing the topology and thermal signals our systems have to react to.

On top of that foundation, three things: provisioning workflows that stand up, update, and decommission customer environments as an application of desired state. The same "declared state" move from ITSF, at multi-customer scale.

A topology controller that continuously reconciles the state of the fleet and exposes it as queryable live status. Rescheduling workflows that respond when a GPU throttles or lands on the wrong side of the interconnect boundary.

All three are built and tested against the sandbox, we fuzz the failure space on demand. Together, they turn managing a fleet of customer environments - which normally grows with the number of customers - into software whose marginal cost per customer trends toward zero.

That's the force multiplier at Dapple scale.

But every reliable multiplier eventually raises the more interesting question:

Where does automation stop, on purpose?

At Dapple we hold customer isolation as a rare decision, not a per-operation checkpoint. In the regulated sectors we target - healthcare, finance, energy - isolation isn't a feature, it's the product.

A machine can guarantee the separation is correct and prove it. What customers pay for is a human to have decided to apply it, and to be able to answer for it. The machine does all the verification work. The human decides and owns it. A handful of times, not on every run.

A human on the hot path doing a machine's work caps throughput. A human at a rare gate who owns a decision does not. This is what my first manager taught me with a phrase I only truly understood years later. Trust does not rule out control.

Automation does all the proof work. It multiplies what a small team can guarantee. The one human gesture we keep changes nothing technical about the result. It only carries responsibility for the decision.

The discipline was never automate everything. It's two instincts held at once. Push the multiplier as far as it goes. Know the single place to stop on purpose. Get both right and they compound.

A force multiplier no one is accountable for isn't leverage, it's exposure.

The automation scales what we can promise.

The boundary is what makes the promise worth trusting.