Skip to main content
| 6 min read

What Network Capacity Delivery Taught Me About Automation

Lessons from leading network capacity automation for ML hardware: model intent explicitly, reserve resources safely, and design for operational reality.

#Network Infrastructure #Automation #Capacity Planning

Earlier in my career at Google, I led cross-functional work on network capacity delivery for machine infrastructure. One part of that work was a switch-port reservation system used to reduce manual planning errors and support new ML hardware launches.

I cannot discuss the internal implementation, but the engineering lessons are not specific to one company. They show up anywhere expensive hardware, long lead times, and multiple operational teams have to meet on the same date.

The main lesson was simple: capacity delivery is not a calculation followed by execution. It is a state-management problem stretched across time.

The spreadsheet was not the real problem

It is easy to blame spreadsheets for manual planning failures. Replacing a spreadsheet with a service, however, does not automatically create a reliable process.

The real difficulty was that several facts changed independently:

  • hardware demand could move between locations or dates;
  • physical installation and network readiness followed different timelines;
  • planned capacity was not the same as installed capacity;
  • a reservation could be valid when created and invalid by execution time;
  • new hardware generations introduced constraints the original process did not know about.

The spreadsheet merely made these disagreements visible. A service could hide them while producing the same bad outcome faster.

Before automating, we needed a better model of the work.

Start with an explicit reservation

A reservation is more than a row saying that a port is taken. It connects an intent to a resource over a period of time.

At minimum, I would expect the model to answer:

  • Who requested the capacity?
  • What resource and quantity are required?
  • Where can the requirement be satisfied?
  • When must the resource become available?
  • Why does the request exist, and what approved plan does it belong to?
  • Which state is planned, reserved, installed, validated, or released?

Those states should not be collapsed into a single ready flag. A resource can be reserved but not installed, installed but not validated, or validated for a requirement that has since changed.

Once the states are explicit, the system can reconcile them instead of guessing what a partially completed workflow means.

Make allocation atomic at the right boundary

Many infrastructure requests need a set of related resources. Returning three quarters of a valid allocation may be worse than returning none, because downstream teams can mistake partial success for deployable capacity.

That does not mean every operation needs one large database transaction. Physical work cannot be made transactional. It means the reservation boundary should match the minimum useful unit of capacity.

The sequence I prefer is:

  1. validate the request and its constraints;
  2. find a complete candidate allocation;
  3. reserve the candidate with a stable request identity;
  4. expose the reservation as pending until downstream validation completes;
  5. either confirm it or release it through an explicit state transition.

Retries should return the existing reservation for the same request, not consume another set of resources.

Keep requirements in data, not branches

New product introduction is where brittle automation becomes obvious. A workflow written around one hardware generation accumulates conditions until nobody is sure which combinations are safe.

The better pattern is to separate the engine from the requirement model. The engine understands concepts such as quantity, compatibility, locality, bandwidth class, redundancy, timing, and validation. A versioned specification describes which of those constraints apply to a particular hardware type.

This does not eliminate code changes. A genuinely new constraint still requires new logic. But it makes the change visible in the model instead of burying another exception inside an allocation function.

Versioning matters here. If a requirement changes after a reservation is made, the system needs to know which version approved that reservation and whether revalidation is required.

Treat time as part of correctness

Most examples of resource allocation ignore time. Capacity delivery cannot.

A resource that exists today may not be available on the required date. A future installation may depend on procurement, power, space, and configuration work that all have different lead times. Two reservations may use the same resource if their windows do not overlap, but only if release dates are trustworthy.

This turns allocation into a planning problem with at least three clocks:

  • the requested need-by date;
  • the estimated readiness date from each dependency;
  • the confidence or risk attached to those estimates.

I learned not to hide uncertainty behind one date. A useful system should show which dependency controls the schedule and how much margin remains. That is more actionable than a red or green status.

Automation does not replace cross-functional decisions

Network capacity delivery crosses hardware planning, network engineering, operations, and the teams consuming the machines. Those groups do not share one planning horizon or one definition of urgency.

The service can enforce invariants and reveal conflicts. It cannot decide every business trade-off.

For example, the system can show that two requests compete for the same constrained capacity and calculate which deadlines are at risk. A person may still need to decide which launch has priority. The automation should preserve that decision as approved intent and execute it consistently, not pretend the prioritization was a technical fact.

This distinction made the design clearer: machines are good at maintaining state, validating constraints, and repeating actions. People remain responsible for decisions that involve changing priorities or accepting risk.

The operational interface is part of the design

The allocation algorithm gets most of the attention, but operators live in the exceptions.

They need to know why a request cannot be satisfied, which constraint failed, what changed since the previous plan, and whether an override will create another conflict. They also need a controlled way to correct bad input, cancel stale intent, and resume work after a dependency recovers.

That led me to judge automation by a different standard. The question is not only, “Did it make the correct allocation?” It is also:

  • Can someone explain the decision later?
  • Can the workflow resume safely after partial failure?
  • Can a new requirement be introduced without invalidating unrelated capacity?
  • Can an operator intervene without creating a second source of truth?

What I carried forward

The switch-port reservation work reduced manual planning errors and helped support zero-downtime launches for new ML hardware. The result mattered, but the reusable part was the model behind it.

Reliable capacity automation needs explicit intent, versioned constraints, idempotent reservations, time-aware state, and an honest boundary between machine decisions and human judgment.

I have carried those ideas into later work on ML and GenAI infrastructure. The resource changes—from network ports to accelerator capacity—but the underlying question remains the same: can the system preserve a complex operational promise while reality keeps changing?

Share Interface:
Return to Index
Vincent's profile photo

Vincent Author

Tech Leader & Architect specializing in LLM Infrastructure, ML Platforms, and Distributed Systems. Passionate about building scalable systems that power the next generation of AI applications.

Related Intelligence

Sep 30, 2025

Seven Automation Rules I Keep Reusing

The practical rules I use when automation has to survive retries, manual intervention, and changing infrastructure requirements.

Read Analysis
Aug 26, 2026

GenAI Capacity Is a Product, Not Just an Accelerator Pool

Lessons from working across ML infrastructure and GenAI serving: reliable capacity requires a product contract across accelerators, entitlements, admission control, scheduling, and operations.

Read Analysis
Aug 27, 2026

From Batch Size to TPU Topology: A Capacity Equation for ML Serving

A practical equation for turning measured batch throughput, latency limits, replica count, and TPU topology into an ML serving capacity estimate.

Read Analysis