Technical Program Manager
Explain an infrastructure program you led end to end or describe a complex infrastructure problem you resolved.
Also asked as: Tell me about an infrastructure or platform initiative you owned from business case through scaled adoption. · Walk me through a complex, cross-organizational infrastructure problem you solved end-to-end. · Describe a process-efficiency or developer-experience program you drove from problem diagnosis to measurable adoption — what made it complex, and how did you take it through to results?
Explain an infrastructure program you led end to end or describe a complex infrastructure problem you resolved.
Situation: At the financial organization, 130+ development teams were losing 60–70% of their time to manual infrastructure provisioning, an 8-person DevOps team had become a hard bottleneck, and the annual cost of inaction was $12.2M. Task: As the TPM, I owned building a self-service Internal Developer Platform end-to-end, and part of that mandate was deciding where machine learning could genuinely cut operational load without taking a judgment call away from a human who needed to make it. Action: I sequenced two AI-assisted capabilities into the platform once the underlying data existed to make them useful, not before: - ML-based anomaly detection, layered on top of the standardized telemetry every service already emitted, so on-call engineers got paged on real deviations instead of static thresholds that fired constant false alarms. - AI-assisted cost-forecasting, layered on top of the cost-center tags the platform injected automatically at provisioning time, generating weekly spend projections so teams could right-size resources proactively instead of reacting to a surprise bill. Both stayed strictly advisory — the anomaly detector paged a human, the forecaster informed a decision — because in a regulated environment I couldn't let a model auto-remediate infrastructure or auto-approve spend. Result: That anomaly detection layer helped the platform hold 99.99% availability, including a sub-15-minute failover during a major AWS regional disruption, and the automated governance plus AI-assisted forecasting together were part of what delivered $2.8M in annual savings — without adding a single person to watch it.