Software Development Manager, Elastic Kubernetes Service (EKS)
Posted Updated
Job Summary
Kubernetes only gets useful once you add things to it: networking, storage, DNS, observability, cert management, AWS resource control. Our team is how Amazon EKS customers get all of that without operating any of it: EKS add-ons and EKS capabilities, installed, upgraded, healed, and security-scanned by us, on live clusters, in every AWS region and partition.
We're hiring a Manager III, Software Development in Seattle to lead that team. You'll own three connected problems: the operator that reconciles add-ons onto customer clusters and repairs them when they drift; the release supply chain that takes an add-on from a partner team's commit to a globally available, CVE-scanned version; and the newest layer: Managed AWS Controllers for Kubernetes (ACK) and kro, where AWS runs the controllers so customers can manage AWS resources from Kubernetes without babysitting the controllers themselves.
- High blast radius, high leverage: an add-on change lands on running production clusters, so correctness and rollout safety are the design constraints, not an afterthought
- You're upstream and downstream of everyone: AWS service teams ship add-ons through your pipeline, and your operator is what actually puts them on customer clusters
- Real open source: ACK and kro are public projects; the team works in the upstream repos, not just around them
Key job responsibilities
- Own the team's charter, roadmap, and delivery for add-on lifecycle, add-on release, and managed ACK/kro capabilities, including the trade-off calls between new capability and paying down operational load
- Hire, grow, and retain SDEs across levels; set the bar for design review, code review, and on-call quality
- Raise release velocity for add-on partner teams while tightening the safety envelope: staged rollouts, bake time, automated validation, fast rollback
- Own operational excellence for a fleet-wide, multi-region, multi-partition service: alarms that mean something, dashboards leadership can read, COEs that produce fixes rather than paragraphs
- Drive the security supply chain for add-on images: scan coverage, time-to-patch, and holding the line with provider teams on CRITICAL/HIGH findings
- Partner with PM and adjacent EKS teams (control plane, connectivity, the internal cluster platform the capability controllers run on) on cross-team designs and launches
- Represent the team in upstream ACK/kro and EKS add-on community work, and in customer and service-team escalations
A day in the life
You start with the health of a system that is always mid-flight: overnight reconciliation and canary signal, an alarm from one region, a partner team's release stuck in a rollout wave. You spend the morning making those either fixed or owned, and asking why the system needed a human at all.
Then the design work. An engineer wants to change how the operator detects drift; another is proposing how a managed controller gets scoped credentials to touch a customer's AWS resources. You review the doc, push on failure modes, and cut scope so it ships in a quarter rather than a year.
Afternoons are the seams between teams. An AWS service team wants their add-on GA'd in a new partition. A container security review needs an answer on patch latency. Your PM wants to know what the roadmap costs if you also take on the next capability type. You have 1:1s, you write, and you fight for the two or three things that actually matter this quarter.
Your customers are both external and internal: EKS customers running these add-ons in production, and the AWS service teams and open-source partners who ship through your pipeline.
About the team
We're the add-ons and capabilities team inside Amazon EKS, based in Seattle. Our mission is simple to state and hard to do: everything a customer adds to their cluster should be installed correctly, upgraded safely, patched fast, and boring to operate.
Culturally we're a code-is-truth team. Designs are written down and argued about; assumptions get checked against the source and against production rather than asserted; the person who finds the sharp edge is expected to file it down, not just report it. On-call is shared and taken seriously, and we treat operational pain as a design bug. We work in public repos alongside upstream maintainers, which keeps the standard honest.
Kubernetes only gets useful once you add things to it: networking, storage, DNS, observability, cert management, AWS resource control. Our team is how Amazon EKS customers get all of that without operating any of it: EKS add-ons and EKS capabilities, installed, upgraded, healed, and security-scanned by us, on live clusters, in every AWS region and partition.
We're hiring a Manager III, Software Development in Seattle to lead that team. You'll own three connected problems: the operator that reconciles add-ons onto customer clusters and repairs them when they drift; the release supply chain that takes an add-on from a partner team's commit to a globally available, CVE-scanned version; and the newest layer: Managed AWS Controllers for Kubernetes (ACK) and kro, where AWS runs the controllers so customers can manage AWS resources from Kubernetes without babysitting the controllers themselves.
- High blast radius, high leverage: an add-on change lands on running production clusters, so correctness and rollout safety are the design constraints, not an afterthought
- You're upstream and downstream of everyone: AWS service teams ship add-ons through your pipeline, and your operator is what actually puts them on customer clusters
- Real open source: ACK and kro are public projects; the team works in the upstream repos, not just around them
Key job responsibilities
- Own the team's charter, roadmap, and delivery for add-on lifecycle, add-on release, and managed ACK/kro capabilities, including the trade-off calls between new capability and paying down operational load
- Hire, grow, and retain SDEs across levels; set the bar for design review, code review, and on-call quality
- Raise release velocity for add-on partner teams while tightening the safety envelope: staged rollouts, bake time, automated validation, fast rollback
- Own operational excellence for a fleet-wide, multi-region, multi-partition service: alarms that mean something, dashboards leadership can read, COEs that produce fixes rather than paragraphs
- Drive the security supply chain for add-on images: scan coverage, time-to-patch, and holding the line with provider teams on CRITICAL/HIGH findings
- Partner with PM and adjacent EKS teams (control plane, connectivity, the internal cluster platform the capability controllers run on) on cross-team designs and launches
- Represent the team in upstream ACK/kro and EKS add-on community work, and in customer and service-team escalations
A day in the life
You start with the health of a system that is always mid-flight: overnight reconciliation and canary signal, an alarm from one region, a partner team's release stuck in a rollout wave. You spend the morning making those either fixed or owned, and asking why the system needed a human at all.
Then the design work. An engineer wants to change how the operator detects drift; another is proposing how a managed controller gets scoped credentials to touch a customer's AWS resources. You review the doc, push on failure modes, and cut scope so it ships in a quarter rather than a year.
Afternoons are the seams between teams. An AWS service team wants their add-on GA'd in a new partition. A container security review needs an answer on patch latency. Your PM wants to know what the roadmap costs if you also take on the next capability type. You have 1:1s, you write, and you fight for the two or three things that actually matter this quarter.
Your customers are both external and internal: EKS customers running these add-ons in production, and the AWS service teams and open-source partners who ship through your pipeline.
About the team
We're the add-ons and capabilities team inside Amazon EKS, based in Seattle. Our mission is simple to state and hard to do: everything a customer adds to their cluster should be installed correctly, upgraded safely, patched fast, and boring to operate.
Culturally we're a code-is-truth team. Designs are written down and argued about; assumptions get checked against the source and against production rather than asserted; the person who finds the sharp edge is expected to file it down, not just report it. On-call is shared and taken seriously, and we treat operational pain as a design bug. We work in public repos alongside upstream maintainers, which keeps the standard honest.