Lead Principal Technical Program Manager

Explain a data center or infrastructure program you led end to end, ideally one involving hardware, technician workflows, or operational tooling across multiple teams.

Also asked as: Walk me through a complex, cross-functional infrastructure program you owned from strategy through general availability. · Describe a program where you owned schedule, risk, and stakeholder alignment across engineering, operations, and vendor teams with no direct authority over any of them.

OwnershipDive DeepBias for ActionDeliver Results

Opening Statement (~60 sec)

I was the Region Build Technical Program Manager who owned three AWS Availability Zone data center launches end-to-end — Dublin, Ireland; Beijing, China; and San Francisco, United States — each starting from a fixed, public-facing launch date and five-plus independently-scheduled workstreams — construction, power, network, security, hardware — with no shared critical path connecting any of them. I owned the program end-to-end, accountable for the schedule, the risk register, and stakeholder alignment across every workstream, with no direct authority over any of the teams actually building it. The strategic bet I made was to stop treating each workstream's schedule as its own truth and instead build one integrated critical-path schedule gated by a binary, evidence-based Day-1 readiness standard, so a slip anywhere showed up immediately as a launch-date risk instead of a surprise two months later. All three AZs launched on their committed dates with zero Sev-1 incidents traced to readiness gaps in the first 30 days — and that same operating model, applied to server fleets, technician workflows, and hardware readiness instead of AZ construction, is the direct playbook I'd bring to OCI's data center server operations.

Situation

1. AWS had committed to fixed, public-facing general-availability dates for new Availability Zones, driven by regional customer demand and data-residency requirements — dates that Commercial and Legal would not move. 2. Delivering an AZ meant coordinating physical construction, utility power interconnection and mechanical commissioning, backbone and metro network builds, security certification, and server hardware racking and capacity provisioning — each owned by a different internal team or external vendor, each running its own schedule with no shared critical path. 3. These workstreams were highly interdependent in ways that weren't visible until they collided — mechanical commissioning couldn't finish without energized power, network validation couldn't start without commissioned space, and hardware capacity provisioning depended on both being done — and each dependency carried its own multi-month lead time. 4. There was no single integrated schedule. Every team could report its own workstream 'on track' while the AZ as a whole was at risk, because no one owned the seams between workstreams.

Task

1. I was assigned as the Region Build TPM with end-to-end ownership of the program — accountable for the launch date, the cross-team schedule, and risk resolution across every dependent workstream, without direct authority over any of the teams building it. 2. My mandate was to partner across Infrastructure, hardware supply chain, network, security, and business stakeholders to build a schedule that held up under real dependency risk, surface and resolve risk before it became a launch blocker, and represent program status credibly to senior leadership. 3. Success meant reaching general availability on the committed date with every Day-1 readiness criterion — power, network, security, hardware capacity — actually met, not just reported green.

Action

  1. 1.Built a single integrated critical-path schedule stitching construction, power/mechanical commissioning, network, security certification, and hardware capacity provisioning into one program plan instead of five independently-tracked workstream schedules.
  2. 2.Ran a structured, program-wide risk register (RAID log) reviewed weekly across every stakeholder team, with each risk assigned a single owner and a resolution date sitting directly on the critical path — not a status color that could sit "yellow" indefinitely.
  3. 3.When utility power interconnection slipped behind the mechanical commissioning window — the single highest-blast-radius risk on the program, I drove a trade-off decision with Data Center Engineering and the utility provider to bring in a temporary generator bridge, protecting the downstream commissioning and hardware timeline while permanent interconnection caught up.
  4. 4.Defined a binary Day-1 launch-readiness gate — power energized and tested, network validated end-to-end, security certification closed, hardware capacity provisioned above forecasted demand — so general availability was earned against evidence, not declared against a calendar date.
  5. 5.Ran the executive reporting cadence myself: a weekly program-health review that led with the critical path and top risks, not a workstream-by-workstream recap, so leadership could see in minutes whether the launch date was actually safe.
  6. 6.Codified the integrated-schedule and readiness-gate model as a reusable playbook, then adapted it for Beijing and San Francisco — each requiring its own highest-risk dependency (in-country regulatory licensing; a constrained urban infill site) — while the core operating model held across all three.

Result

1. All three Availability Zones — Dublin, Beijing, and San Franciscolaunched on their committed general-availability dates, with every Day-1 readiness criterion met at go-live, not retrofitted after. 2. Dublin's power-interconnection risk was resolved without moving the launch date, protecting a downstream commitment already made to customers. 3. Zero Sev-1 operational incidents traced to launch-readiness gaps in the first 30 days post-GA across the three AZs — the direct result of gating GA on evidence instead of a date. 4. The integrated-schedule and readiness-gate model became the reusable pattern for every subsequent AZ, cutting the time to stand up a credible cross-team program schedule since the framework — not just the lessons — transferred.

Closing Statement (~60 sec)

What made this a program-management problem, not a construction problem, was that the real risk never lived inside any single workstream — it lived in the seams between construction, power, network, security, and hardware, where no team had visibility past its own boundary. That's the same pattern I'd expect to run at OCI's scale, just with server operations, technician workflows, and hardware support replacing AZ construction as the workstreams: the hardest failures aren't inside any one team's system, they're in the dependency and ownership seams between task execution, inventory coordination, guided procedures, and the platforms that support them. The operating model I'd bring is the same one I proved on three AZs — one integrated critical path instead of several independent ones, risk owned and retired before it hits the path, and readiness gated on evidence — the same discipline that underpins the governance mechanisms, KPIs, and launch-readiness frameworks this role is scoped to establish. The lesson I carry forward is that at this scale, the program's job isn't managing any one team's execution — it's making the seams between teams visible and owned before they become the incident (a RAID log, for anyone unfamiliar, is a shared tracker of Risks, Assumptions, Issues, and Dependencies, reviewed on a fixed cadence with a single named owner per line).