Technical Program Manager

How do you motivate others and build partnerships across teams to achieve a collective goal?

Also asked as: Tell me about a time you motivated a team or organization to embrace a difficult, unpopular change. · Describe how you built a partnership or coalition across teams with competing priorities to achieve a shared outcome. · How did you identify and troubleshoot project bottlenecks by building relationships across technical and non-technical teams?

Earn TrustOwnershipAre Right, A LotDeliver Results

Opening Statement (~60 sec)

I want to tell you about a program where I inherited a 4-year organizational stall and turned it around in 16 months — but the hardest part wasn't technical. At a large financial institution, I led the AWS Small Account (Gen3) Migration, moving 2,114 Business Applications out of a handful of massive, shared legacy AWS accounts — some housing up to 246 unrelated applications — into isolated, dedicated small accounts. After four years, adoption sat below 10%, because no single team could compel 11 independent business divisions, each with its own roadmap and budget, to collectively own a risk none of them wanted. I didn't lead with a mandate. I led with three weeks of listening, then rebuilt the program around making an abstract cyber risk personally real to each division, a partnership model that gave Cybersecurity, our CORE infrastructure team, and AWS ProServe a genuine stake instead of a review role, and an $11M shared-investment pitch that turned 'you must migrate' into 'you're not migrating alone.' The result: within 16 months — two months ahead of the 2026 deadline — all 2,114 applications reached 100% adoption, blast radius risk fell 60%+, compliance rose 90%, and the partnership model I built was adopted as the standard for three other transformation programs.

Situation

1. The financial institution had consolidated over 2,100 Business Applications into a handful of large, shared legacy AWS accounts, with Gen1 accounts averaging 77 co-tenanted applications (up to 246) and Gen2 averaging 35 (up to 126). 2. This concentration created three compounding, quantified risks Cybersecurity had escalated to the C-suite: Blast Radius (a single incident could cascade across dozens of unrelated applications), Lateral Movement (a threat actor compromising one BA could 'role shop' or 'security-group shop' to reach unrelated systems in the same account), and Operational Complexity (IP exhaustion, configuration drift, escalating maintenance). 3. A formal migration initiative had been running for nearly four years with adoption of the safer Gen3 'Small Account' model still below 10% — not because divisions disagreed with the goal, but because no one owned it: priorities kept losing to product roadmaps, the existing tool (Pathfinder) required 3–6 months of manual effort per application, and there was no funding or governance structure to absorb that cost. 4. Three root causes explained the stall: fluctuating business-unit priorities, no clear ownership or executive mandate, and a hard tooling gap that made every migration a slow, inconsistent, manual undertaking.

Task

1. I was brought in as the TPM to relaunch this as an enterprise-wide, time-bound program — drive 100% adoption across 11 Lines of Business, migrating all 2,100+ applications by the 2026 deadline, entirely through influence, not authority: none of the 11 divisions reported to me, and I inherited no pre-approved budget or governance foundation. 2. My mandate had two inseparable halves: motivate 11 competing business units — each with its own roadmap, risk tolerance, and engineering leadership — to align on a shared deadline; and build genuine partnerships, not just coordination, with Cybersecurity, CORE Infrastructure, and AWS ProServe so the program had technical and financial backing strong enough to remove the bandwidth objection that had stalled it for four years. 3. Success meant securing executive sponsorship and $11M in cross-divisional funding for a central migration factory and automation tooling, cutting migration time per application by 70%, and converting every unmigrated application into a formally owned, documented risk rather than an invisible delay.

Action

  1. 1.Ran three weeks of structured discovery workshops across application teams, divisional TPMs, Cybersecurity, and CORE engineers before proposing anything, asking only two questions — 'What is making this hard?' and 'What would you need to say yes?' — which surfaced that teams weren't resisting the goal, they were resisting absorbing all the migration pain with zero budget, tooling, or governance support.
  2. 2.Partnered with Cybersecurity to run formal threat-model assessments across Gen1/Gen2 accounts, then personally translated the findings from security abstractions ('lateral movement', 'blast radius') into plain-language, per-application risk inventories — opening CIO conversations with 'if this application were compromised tomorrow, here are the other 40 applications in the same account also at risk,' which reframed the conversation from 'when must we' to 'how fast can we move.'
  3. 3.Built a tailored value proposition per audience from the same underlying data: 70% migration-time reduction and dedicated factory engineers for engineering leads, a per-LOB ROI model showing the $11M investment would avoid an estimated $120M in future costs for finance partners, and organizational-resilience/reputational-risk framing for executives.
  4. 4.Built a genuine partnership coalition rather than a review pipeline: co-developed a Risk Sloping prioritization methodology with Cybersecurity that they owned and trusted, worked with CORE to define migration-complexity ratings and three target-state architecture patterns, and co-designed a two-phase acceleration plan with AWS ProServe — Pathfinder automation for S3/RDS/DynamoDB migrations, then a central migration factory with dedicated TPMs and architects aligned to each LOB.
  5. 5.Structured the $11M investment as a per-LOB contribution proportional to each division's application count, then ran alignment sessions with each LOB's finance team to negotiate cost allocation — pitching it as 'you are contributing to a shared engine that does the heavy lifting for all of us,' not a tax on their budget.
  6. 6.Resolved the core conflict — Cybersecurity's urgency versus application teams' committed roadmaps, the actual driver of the four-year stall — with three levers: risk-based wave sequencing so the highest-risk applications moved first while lower-risk teams kept flexible timelines; the funded migration factory so no team migrated alone; and formal risk registration in FUSE so any deferral became a documented, owned risk instead of an invisible delay.
  7. 7.Secured a signed memorandum from our AE to all divisional CIOs making migration a C-level, 18-month non-negotiable directive, then built the delivery infrastructure to execute at speed — a 9-wave migration strategy tracked via ETBs, CI/CD enforcement blocking new components in legacy accounts, CloudHydra self-service provisioning, and CloudRadar dashboards giving every partner the same real-time source of truth.

Result

1. Within 16 months — two months ahead of the 2026 deadline — all 2,114 applications reached 100% adoption of the Gen3 small-account model, up from below 10% after four years. 2. Blast radius risk fell 60%+ (to under 30% within the first 12 months) and lateral movement risk was eliminated within legacy accounts; compliance posture rose 90%. 3. Automation cut migration time per application by 70%, and CI/CD enforcement blocked 5,459 new legacy components, avoiding $21M in future migration costs and 218,360 engineering hours. 4. The $11M shared investment avoided an estimated $120M in projected labor and cloud costs, and ongoing maintenance costs fell ~25%. 5. All 11 business divisions signed onto one program schedule; the Risk Sloping methodology became the enterprise standard, and the migration-factory partnership model was adopted by three other business units — the biggest win wasn't the metrics, it was that 11 divisions went from avoiding a shared problem to co-owning a shared solution.

Closing Statement (~60 sec)

What made this a motivation-and-partnership problem rather than a pure technical migration was that no authority I had could force 11 independent divisions to reprioritize their roadmaps — the unlock only came from making an abstract risk personally real to each leader, and from treating Cybersecurity, CORE, and AWS ProServe as co-owners with a real stake rather than reviewers I had to get past. The pattern I'd bring here is the same one that broke a four-year stall in 16 months: listen before proposing, translate the same data into the currency each partner already cares about, and design the ask so saying yes is easier than saying no — because at enterprise scale, people don't resist change, they resist being asked to absorb its cost alone.

GenAI wasn't a side project bolted onto this migration — it was the mechanism that made a 70% cut in per-application migration time achievable at 2,100+ application scale. This tab covers the business case for using AI as an enabler, the multi-agent architecture behind it, how it was implemented end-to-end, and the specific risks of running AI-driven automation inside a regulated banking environment.

AI Business Case

DimensionDetail
The problem with the manual baselinePathfinder, our existing migration tool, had no end-to-end automation — every migration took 3–6 months of manual discovery, mapping, and cutover work. At 2,100+ applications, that baseline implied a 36–48+ month, 50+ specialist effort that no linear increase in headcount could compress into an 18-month regulatory window.
Why GenAI specifically fit this problemThe highest-effort tasks — dependency discovery, IaC generation, migration runbook authoring, compliance mapping, cost right-sizing — are structured, pattern-matching problems grounded in existing metadata (CMDB, AWS Application Discovery Service, RVTools), not novel judgment calls. That made them a strong automation target for LLM-based agents rather than a speculative AI use case.
The investment caseGenAI automation was proposed as part of the $11M tooling and migration-factory investment (with AWS ProServe), projected to cut migration time 70%, reduce LOB effort 20%, and avoid an estimated $120M in future labor and cloud costs — structured with a GenAI proof-of-concept and a go/no-go gate before committing to production automation.
Where it created leverage beyond costAutomating the repetitive majority of the work freed scarce SME architects and TPMs to focus on the highest-risk, highest-complexity applications in the pilot wave, instead of spreading fixed human capacity evenly across all 2,100+ applications regardless of risk.
Business outcome deliveredAI-assisted automation contributed directly to the program's headline results: the 70% migration-time reduction, $4.2M saved in migration costs (73% effort reduction), and a further $2.3M in annual savings identified through AI-driven right-sizing.

AI Architecture

LayerComponentRole
User InterfaceNatural-language chat interface + executive dashboardsLets both engineers and business stakeholders query migration status, view compliance reports, and adjust plans in plain language.
AI Agent OrchestrationSupervisor AgentActs as the central brain — delegates tasks to specialized agents (Dependency Mapping, Wave Planning, Resource Optimization, Compliance Validation, Cost Efficiency) and assembles their outputs into a single recommendation.
AI Foundation ServicesAmazon Bedrock + Claude 3.5 Sonnet, Amazon OpenSearchProvides the reasoning engine for complex judgment calls (e.g., rehost vs. replatform) and a vector database for Retrieval-Augmented Generation against curated migration best-practice Knowledge Bases.
Data & IntegrationAWS Application Discovery Service (ADS), CMDB, RVToolsBridges on-premises legacy inventory data into the platform so every agent decision is grounded in real, current infrastructure state rather than model assumptions.
Target EnvironmentAWS Organizations, Transit GatewayThe destination multi-account, hub-and-spoke architecture the agents provision and migrate workloads into.

What was the end-to-end AI workflow, from discovery to cutover?

Discovery: the Supervisor Agent pulls real-time inventory from AWS Application Discovery Service so no decision is based on stale or hallucinated infrastructure state. Reasoning: the Dependency Mapping Agent passes application logs to Amazon Bedrock/Claude 3.5 Sonnet to identify application boundaries and "chatty" groups that must migrate together. Planning: the Wave Planning Agent sequences the resulting groups into risk-based waves. Approval: the plan is presented to a human for sign-off before any migration tool is triggered. Governance: the Compliance Validation Agent audits the plan against PCI-DSS/SOC2/internal frameworks before the user ever sees the recommended strategy. Execution: AWS Step Functions and AWS Application Migration Service (MGN) carry out the approved plan, with Amazon EventBridge streaming progress back to the Supervisor Agent for the live dashboard.

AI Implementation

CapabilityImplementation Detail
Orchestration engineAWS Step Functions ran the migration as a state machine, so a failed step (e.g., a data-sync timeout) triggered automated retry or rollback rather than manual intervention.
Automation glueAWS Lambda handled account-readiness checks against AWS Organizations and passed optimized instance specs from the Resource Optimization Agent into the execution pipeline.
Replication engineAWS Application Migration Service (MGN) performed block-level replication; Step Functions gated cutover on ReplicationStatus reaching full consistency.
Post-launch configurationAWS Systems Manager applied monitoring agents and security hardening automatically once an instance booted in the target account.
Data pipelineIngestion from RVTools/CMDB, processing via Kinesis and AWS Glue, storage across S3/DynamoDB/OpenSearch, and visualization in Amazon QuickSight.
Measured automation outputResolved 2,300+ dependency conflicts pre-migration, auto-generated 2,100+ application runbooks and IaC templates for 200+ target accounts, and reached a 94% first-time migration success rate with 99.97% availability and zero data loss.

AI Risks

RiskDescriptionMitigation
Hallucinated or stale infrastructure stateAn agent recommending a migration plan based on inaccurate assumptions about what actually existed in a legacy account could produce a materially wrong plan.Every agent decision was grounded in real-time data pulled directly from AWS Application Discovery Service and CMDB rather than model-generated assumptions, with RAG against curated, versioned migration best-practice Knowledge Bases.
Automation bypassing regulatory reviewIn a regulated banking environment, a fully automated pipeline that skipped human review could create compliance exposure the organization couldn't defend to an examiner.Every plan required human-in-the-loop approval before any migration tool executed, the Compliance Validation Agent audited every plan against PCI-DSS/SOC2/internal frameworks before the user saw it, and every action was logged immutably to CloudTrail for audit.
Cascading automated failuresAn automated rollback or remediation acting on a systemic error could propagate across many applications simultaneously before a human noticed.A circuit-breaker pattern paused all executions within an affected dependency ("Affinity Group") cluster on repeated failures, retries used exponential backoff for transient errors, and non-transient failures triggered compensating-transaction rollback (infrastructure teardown, MGN state reset, DNS/Transit Gateway reversion).
Over-privileged or exposed access during migrationTemporary elevated access needed for migration tooling could be exploited or left in place after cutover.A temporary, narrowly scoped "Migration Role" was used only during replication, and the Compliance Agent automatically downgraded it to a least-privilege "Production" role — with KMS re-encryption to the target account's keys — immediately once migration validation completed.
Adoption and trust risk among SMEsMigration architects and application teams could distrust AI-generated wave plans and IaC in a zero-downtime, regulated environment, quietly reverting to manual work and eroding the automation's value.Validated the approach on a 150–250 application pilot wave before scaling, kept every recommendation subject to human sign-off, and trained 120+ staff on AI-assisted methodology so the tooling visibly augmented — rather than replaced — SME judgment.