Technical Program Manager

Explain an infrastructure program you led end to end or describe a complex infrastructure problem you resolved.

Also asked as: Tell me about an infrastructure or platform initiative you owned from business case through scaled adoption. · Walk me through a complex, cross-organizational infrastructure problem you solved end-to-end. · Describe a process-efficiency or developer-experience program you drove from problem diagnosis to measurable adoption — what made it complex, and how did you take it through to results?

OwnershipEarn TrustInvent and SimplifyAre Right, A LotDeliver Results

Opening Statement (~60 sec)

I led the design and delivery of the financial organization's Internal Developer Platform (IDP) — a self-service layer, fronted by a developer portal, that replaced weeks of manual infrastructure provisioning with a standardized golden path to production. The trigger was a measured productivity crisis: time-motion studies across our development organization showed engineers losing the majority of their time to non-coding work [environment setup, infrastructure provisioning, and navigating compliance gates], with an overloaded DevOps team unable to keep pace as the single point of provisioning. As the TPM, I owned this end-to-end — from planning and cross-org alignment through execution, governance, delivery, and adoption. I built the bottom-up financial model that secured VP-level buy-in across nine organizations, made the strategic call to build platform-as-product — treating the platform itself as a living product with its own roadmap and long-term ownership, rather than a one-time tooling project — sequenced a three-phase roadmap that delivered value every few months instead of asking the business to wait 18 months, and governed four concurrent engineering workstreams through to full delivery — time to first deployment dropped from 6 weeks to 2 hours, and the platform reached 90% organic adoption in 8 months, without a single top-down mandate.

Situation

1. At the financial organization, developer productivity had become a quantified crisis, not just an anecdotal complaint. Time-motion studies — a structured observational technique borrowed from industrial engineering that tracks exactly how workers allocate their time — across 130+ development teams showed developers spending 25 hours per week (60–70% of their time) on non-coding work: environment setup, infrastructure provisioning, and navigating compliance gates. 2. Our 8-person DevOps team had become a hard organizational bottleneck, with 4–6 week provisioning lead times for anything a team needed to stand up. 3. Developer satisfaction had fallen to 5.2 out of 10, and feature delivery cycles had stretched from 2 weeks to 8 weeks as teams absorbed the provisioning and compliance overhead themselves. 4. The annual cost of inaction — productivity loss, security incidents, and developer turnover combined — was $12.2M, and none of it was being addressed by adding more DevOps headcount, since that cost scaled linearly with the problem rather than solving it.

Task

1. I had to design a strategic solution end-to-end — not deploying a tool, but quantifying the problem, building a defensible business case, gaining VP-level buy-in across nine organizations with competing priorities, and sequencing a roadmap that delivered measurable value in months, not after an 18-month wait. 2. I framed this from day one as building a platform-as-product — an Internal Developer Platform (IDP): a curated, self-service layer of tools, APIs, and automated workflows, fronted by a developer portal, that abstracts infrastructure, security, and compliance complexity behind a standardized 'golden path' to production. 3. I covered program strategy, the financial model, cross-org stakeholder alignment, the technical architecture bets, and delivery — owning the gap between 'we have a productivity problem' and a production platform that teams adopted by choice.

Action

  1. 1.Built the financial case bottom-up from four measurable cost categories — developer productivity loss ($10.4M), DevOps bottleneck costs ($800K), security incidents ($600K), and developer turnover ($400K) — totaling $12.2M in annual cost of inaction, then projected $3.16M in annual value with an 18-month payback period, and translated the same numbers into different language for each audience: 'focus on code, not infrastructure' for developers, 'shift-left with controls that can't be bypassed' for security, 'accelerate time-to-market and reduce operational risk' for business leaders.
  2. 2.Used a RACI model to formalize cross-organizational ownership across the nine participating orgs, and stood up a Developer Experience Council of lead engineers from each business unit so the platform was built *with* developers rather than mandated *to* them.
  3. 3.Converted Internal Audit from gatekeeper to co-owner by inviting them into the program's Definition of Done and building them a real-time compliance dashboard with kill-switch visibility — the ability to see every deployment's compliance status in real time and manually halt any non-compliant one before it reaches production — replacing 10-day manual approval tickets with automated policy enforcement.
  4. 4.Sequenced a risk-validated, three-phase roadmap instead of a single 18-month delivery: Phase 1 (months 1–6) proved speed-to-value with basic self-service provisioning and standardized CI/CD templates; Phase 2 (months 7–10) converted compliance from a manual tollgate into automated, Policy-as-Code guardrails; Phase 3 (months 11–18) scaled the platform into a full self-service product with a Backstage developer portal, advanced observability, and cost governance — each phase deliberately sequenced because it depended on capabilities the prior phase established.
  5. 5.Resolved a structural resource conflict between two VP-level escalations competing for the same platform engineering capacity — a Retail Banking VP's request for real-time fraud-detection infrastructure and an Auto Finance VP's request to accelerate a regulatory-driven Dealer Portal migration — using a Cost of Delay / CD3 framework and a weighted scoring rubric (Regulatory Risk 40%, Revenue Impact 35%, Architectural Leverage 25%) that produced a data-backed sequencing recommendation both VPs accepted.
  6. 6.Organized the build into four concurrent engineering workstreams — Platform Infrastructure & Provisioning, Developer Portal & Golden Path Tooling, Security & Policy Automation, and Observability & Cost Governance — each with a dedicated technical lead and an interface contract requiring 48-hour notice before any change that could affect another workstream.
  7. 7.Managed the program's two highest risks — a new architectural single point of failure introduced by centralizing provisioning, and 'Shadow IT' adoption risk in a 'you build it, you own it' culture — and governed scope creep through a formal Change Control Board, including converting a mid-Phase-2 security-scanning request into a temporary compliance shim with a negotiated technical-debt repayment agreement rather than letting it derail the milestone.

Result

1. Time to first deployment dropped from 6 weeks to 2 hours, and feature delivery cycles returned from an 8-week cadence to 2 weeks. 2. Developer satisfaction rose from 5.2 to 8.4 out of 10, consistent with industry data showing IDP investment and enabling-team independence drive measurable productivity gains. 3. Security vulnerabilities in production fell 85% through automated policy enforcement, zero policy violations reached production post-enforcement, and quarterly audit preparation time dropped 70%. 4. The platform sustained 99.99% availability, including failing over a major AWS regional disruption in under 15 minutes, and delivered $2.8M in annual cost savings through standardized cloud resource usage. 5. 90% organic adoption within 8 months — without a single mandate — validating the core thesis: durable platforms are earned through stakeholder co-ownership and disciplined scope governance, not imposed top-down.

Closing Statement (~60 sec)

What made this a program-management problem rather than a pure infrastructure build was that the sequencing, the cross-org trade-offs, and the adoption model all had to be decided before the first Crossplane composition shipped — and none of those decisions had a clean 'right' answer, only a better trade-off given the constraint. Platform engineers could tell me whether a Composition API was technically sound; they couldn't tell me whether to sequence the Dealer Portal ahead of fraud detection, or how to convert a skeptical DevOps team into a co-owner instead of a blocker. That's the gap I closed: I turned a productivity crisis into a quantified business case, a scoring framework that resolved VP-level conflicts with data instead of authority, and an adoption model — Platform Champions, a compliance tax instead of a mandate — that earned 90% organic adoption. The pattern I'd bring here is the same: quantify the cost of the status quo before proposing a fix, sequence the riskiest, highest-leverage architecture first, and design adoption to be the rational economic choice, not a directive. (A Crossplane composition is a reusable, declarative infrastructure blueprint — it maps a developer's simple configuration request to a fully provisioned, security-configured set of cloud resources, without anyone touching a cloud console directly.)

AI showed up in this platform in two advisory capacities, not as an autopilot: ML-based anomaly detection layered onto the observability data every service already emitted, and AI-assisted cost forecasting layered onto the cost-attribution tags injected automatically at provisioning time.

Explain an infrastructure program you led end to end or describe a complex infrastructure problem you resolved.

Situation: At the financial organization, 130+ development teams were losing 60–70% of their time to manual infrastructure provisioning, an 8-person DevOps team had become a hard bottleneck, and the annual cost of inaction was $12.2M. Task: As the TPM, I owned building a self-service Internal Developer Platform end-to-end, and part of that mandate was deciding where machine learning could genuinely cut operational load without taking a judgment call away from a human who needed to make it. Action: I sequenced two AI-assisted capabilities into the platform once the underlying data existed to make them useful, not before: - ML-based anomaly detection, layered on top of the standardized telemetry every service already emitted, so on-call engineers got paged on real deviations instead of static thresholds that fired constant false alarms. - AI-assisted cost-forecasting, layered on top of the cost-center tags the platform injected automatically at provisioning time, generating weekly spend projections so teams could right-size resources proactively instead of reacting to a surprise bill. Both stayed strictly advisory — the anomaly detector paged a human, the forecaster informed a decision — because in a regulated environment I couldn't let a model auto-remediate infrastructure or auto-approve spend. Result: That anomaly detection layer helped the platform hold 99.99% availability, including a sub-15-minute failover during a major AWS regional disruption, and the automated governance plus AI-assisted forecasting together were part of what delivered $2.8M in annual savings — without adding a single person to watch it.