Technical Program Manager

Explain a data center build that you led end to end at AWS.

Also asked as: Walk me through an AWS Availability Zone launch you owned from architecture through go-live. · Tell me about a complex, cross-functional infrastructure program you led end to end. · Describe an infrastructure build where you owned schedule, risk, and stakeholder alignment across multiple engineering and construction teams.

OwnershipDive DeepBias for ActionAre Right, A LotDeliver Results

Opening Statement (~60 sec)

At AWS, I was a Region Build Technical Program Manager, and I owned three Availability Zone data center launches end-to-end — Beijing, China; Dublin, Ireland; and San Francisco, United States — each from initial site engineering through general availability. I'll walk through Dublin, since it's the clearest example of the full complexity: I was accountable for sequencing construction, power and mechanical commissioning, backbone and fiber networking, security and compliance certification, and capacity provisioning across more than a dozen internal AWS teams and external vendors, against a fixed launch date none of those teams individually controlled, with no shared schedule connecting any of them.

The operating model I built there — one integrated critical-path schedule, risk owned and retired before it hit the path, and a launch gated on measurable readiness rather than a date on a calendar — is the same model I reused on Beijing and San Francisco, and it's the direct analog to what Meta's Core Infrastructure org needs to run compute, storage, networking, and platform programs across a fleet an order of magnitude more heterogeneous.

Situation

1. AWS had committed to a fixed general-availability date for a new Availability Zone in Dublin, Ireland, driven by regional customer demand and data-residency requirements — the date was public-facing and not something Commercial or Legal would move.

2. Delivering an AZ meant coordinating physical construction and fit-out, power utility interconnection and mechanical (cooling) commissioning, backbone and metro fiber builds, security hardening and compliance certification, and server/network hardware racking and capacity provisioning — each owned by a different internal team or external vendor, each running its own schedule with no shared critical path.

3. These workstreams were highly interdependent in ways that weren't visible until they collided: mechanical commissioning couldn't complete without utility power energized; network backbone couldn't be validated without mechanical spaces ready for hardware; security certification couldn't close until both power and network were live — and each dependency carried its own multi-month lead time.

4. There was no single integrated schedule. Every team could report its own workstream as 'on track' while the AZ as a whole was at risk, because nobody owned the seams between workstreams.

Task

1. I was assigned as the Region Build TPM with end-to-end ownership of the Dublin AZ program — accountable for the launch date, the cross-team schedule, and risk resolution across every dependent workstream, with no direct authority over any of the teams building it.

2. My mandate mirrored what AWS asks of this role generally: partner with Infrastructure, Service Teams, and Business Development to design an efficient build, build a schedule that held up under real dependency risk, assess and resolve risk before it became a launch blocker, and represent program status credibly to senior leadership.

3. Success meant Dublin reaching general availability on the committed date, with every Day-1 readiness criterion — power, network, security, capacity — actually met, not just reported green.

Action

  1. 1.Built a single, integrated critical-path schedule that stitched together construction, power/mechanical commissioning, backbone networking, security certification, and capacity provisioning into one program plan — instead of five independently-tracked workstream schedules — so a slip in one workstream showed up immediately as a launch-date risk, not a surprise two months later.
  2. 2.Ran a structured, program-wide risk register (RAID log) reviewed weekly across every stakeholder team, with each risk assigned a single owner and a resolution date on the critical path — not a status color that could sit "yellow" indefinitely.
  3. 3.When utility power interconnection slipped behind the mechanical commissioning window — the single highest-blast-radius risk on the program, since every downstream workstream depended on energized powerI drove a trade-off decision with Data Center Engineering and the utility provider to bring in a temporary generator bridge, protecting the mechanical and network commissioning timeline while permanent interconnection caught up, rather than letting the whole program slip to the slowest dependency.
  4. 4.Defined a binary Day-1 launch-readiness gate — power energized and tested, network backbone validated end-to-end, security certification closed, capacity buffer provisioned above forecasted demand — so general availability was earned against evidence, not declared against a calendar date.
  5. 5.Ran the executive reporting cadence myself: a weekly program-health review to senior leadership that led with the critical path and top risks, not a workstream-by-workstream status recap, so a VP could see in minutes whether the launch date was actually safe.
  6. 6.Codified the integrated-schedule and risk-gate model as a reusable playbook, then adapted it for Beijing and San Francisco — Beijing required layering in a joint-venture operating structure and in-country regulatory/licensing dependencies that Dublin didn't have; San Francisco required resequencing around a constrained urban infill site and tighter fiber-path options. The core operating model held across all three; what changed each time was which dependency carried the highest risk.

Result

1. All three Availability Zones — Dublin, Beijing, and San Franciscolaunched on their committed general-availability dates, with every Day-1 readiness criterion met at go-live, not retrofitted after.

2. Dublin's power-interconnection risk was resolved without moving the launch date, protecting a downstream commitment to customers who had already been told the AZ would be available.

3. Zero Sev-1 operational incidents traced to launch-readiness gaps in the first 30 days post-GA across the three AZs — the direct result of gating GA on evidence instead of a date.

4. The integrated-schedule and risk-gate model I built for Dublin became the reusable pattern for Beijing and San Francisco, cutting the time to stand up a credible cross-team program schedule on each subsequent AZ, since the framework — not just the lessons — transferred.

Closing Statement (~60 sec)

What made this a program-management problem, not a construction problem, was that the real risk never lived inside any single workstream — it lived in the seams between construction, power, networking, security, and capacity, where no team had visibility past its own boundary.

That's precisely the pattern Meta's Core Infrastructure org is running at a much larger scale: millions of servers, multiple GPU architectures, thousands of interconnected services, and a multi-cloud footprint across AWS, Oracle, Google Cloud, and emerging neoclouds — where the hardest failures aren't inside one team's system, they're in the dependency and capacity seams between compute, storage, networking, and the platforms sitting on top of them. The operating model I'd bring is the same one I proved on three AZs: one integrated critical path instead of five independent ones, risk owned and retired before it hits the path, and readiness gated on evidence — the same discipline that underpins the kind of SLO-driven reliability accountability and change-safety programs Meta is scaling today.

The lesson I carry forward is that at this scale, the program's job isn't managing any one team's execution — it's making the seams between teams visible and owned before they become the incident.

Follow-Up Questions

Meta's Core Infrastructure org is delivering next-generation service management platforms to power AI inference at scale — model orchestration, deployment lifecycle, and a unified control plane across thousands of interconnected services. The dependency-sequencing and readiness-gating discipline from AZ builds maps directly onto standing up AI infrastructure capacity, but the workload profile changes what "readiness" actually means.

Tell me about your direct, hands-on experience with AI infrastructure during your data center builds — including any pivot from CPU-first to GPU-first design, and the specific GPU-first workloads you built capacity for.

This is a pivot I lived through mid-build, not a slide I've presented — it changed the physical design of an AZ I was actively delivering.

The Trigger: Midway through the San Francisco AZ build, the capacity plan I'd already baselined against general-purpose compute stopped matching reality — AI/ML training and inference demand was accelerating faster than the original forecast, and I had to re-sequence the build around a GPU-first design without moving the committed launch date.

The Approach: I treated the gap as a formally tracked risk, not just a scope change, and re-sequenced through a phased split rather than a single big-bang redesign — a funded Phase 1, the highest-confidence GPU demand with committed service-team sign-off, absorbed inside a re-densified subset of the existing power and cooling envelope before GA, and a Phase 2 post-GA expansion for lower-confidence, longer-lead demand on its own schedule. That phased split — detailed in the Risks tab — is what let "GPU-first design" and "unchanged launch date" both hold true at once.

Power & Cooling: Power density assumptions that held for CPU racks didn't hold for GPU racks — I re-baselined the power and cooling plan around materially higher power draw per rack, and for the highest-density zones, shifted from air cooling to direct liquid cooling, reopening decisions with Data Center Engineering I'd already treated as closed.

Network Fabric: General-purpose compute ran fine on standard leaf-spine, but GPU clusters needed a low-latency, high-bandwidth interconnect — AWS's Elastic Fabric Adapter (EFA) — validated end-to-end before I'd call a cluster ready, the same evidence-based readiness gate I used for power and network everywhere else on the AZ, just with a different bar for "done."

GPU-First Workloads: I provisioned EC2 P4d (NVIDIA A100) and later P5 (NVIDIA H100) capacity into UltraCluster-scale pools for distributed model training, alongside a growing footprint of AWS Trainium and Inferentia — Amazon's own custom AI silicon — running inference for high-throughput recommendation and generative workloads.

The Complexity: This AZ wasn't provisioning one kind of capacity — it was provisioning at least three distinct hardware profiles (general-purpose CPU, NVIDIA GPU, and custom AI silicon), each with its own power, cooling, and readiness criteria, inside the same program and the same launch date.

That's the exact pattern I'd expect to run again at Meta, just wider: a heterogeneous fleet across multiple GPU architectures, where the program's job isn't picking one hardware profile and optimizing for it — it's building a capacity and readiness model flexible enough to absorb a mid-build pivot like the one I lived through in San Francisco, without the launch date, the budget, or the reliability bar moving. I didn't learn that discipline from a case study; I re-baselined a live build around it.

Leaf-spine: a two-tier data center network design where every 'leaf' (top-of-rack) switch connects directly to every 'spine' switch, giving any server a predictable, low-hop path to any other server. It's the standard topology for general-purpose compute, but it doesn't provide the uniform, high-bandwidth, all-to-all connectivity that distributed GPU training clusters need — which is why that workload calls for a purpose-built fabric like EFA instead.

How does building physical AZ capacity for general compute differ from building capacity for AI training and inference workloads, and how would that change your Region Build playbook?

The critical-path discipline stays the same; the highest-risk dependency moves. A general-purpose AZ's riskiest dependency is usually utility power interconnection and mechanical commissioning at a fairly uniform power density per rack. An AI training or inference build changes that in three ways: power density per rack runs several times higher, which usually forces liquid or direct-to-chip cooling instead of standard air cooling; the network fabric requirement shifts from a general leaf-spine design to a low-latency, high-bandwidth interconnect built for distributed training or large-batch inference; and capacity provisioning stops being just 'racks and power' and becomes 'validated cluster topology' — a set of nodes isn't ready until the interconnect fabric between them has been proven end-to-end, not just powered on. I'd keep the same integrated-schedule and evidence-based readiness gate I used in Dublin, but the gate criteria and the highest-blast-radius risk owner both change.

Meta talks about making infrastructure "fully operable by AI agents" — intent-driven, agent-executed workflows. Where would a program built on Day-1 readiness gates and RAID logs fit into that shift?

Structured, evidence-based readiness gates and a disciplined risk register are exactly the kind of deterministic, auditable process that's a strong candidate for agent-executed automation — the criteria are explicit, the evidence is checkable, and the ownership is already assigned. In that model, my job as the program owner shifts from manually chasing status updates to defining the intent clearly enough for an agent to execute against — what "ready" means, what evidence a risk resolution requires — and letting agents run the checks and surface exceptions, while I keep the judgment calls: blast-radius trade-offs, like the power-bridge decision in Dublin, that carry real cost and shouldn't be made autonomously.

With thousands of interconnected services across a heterogeneous fleet, how do you keep a dependency map like the one you built for Dublin from becoming unmanageable?

The same principle that worked at AZ scale — centralize on the handful of dependencies with the highest blast radius, and don't try to personally track everything else — has to scale through delegation and contracts rather than a bigger spreadsheet. At Dublin, I owned the seams between five workstreams myself because five was tractable. At thousands of interconnected services, that only works if most service-to-service dependencies are governed by a shared SLO contract the owning teams manage themselves, and the program level only tracks the small number of shared, high-blast-radius dependencies — shared control planes, capacity pools, core network paths — the way I tracked power, network, and security at the AZ level. Scale changes the tooling; it doesn't change the principle of finding the few dependencies that can take the whole program down and owning those explicitly.