Senior Cloud / AWS Infrastructure Engineer
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Cloud / AWS Infrastructure Engineer based in the United States.
As a Senior Cloud / AWS Infrastructure Engineer, you’ll build, operate, secure, and evolve a large-scale AWS environment supporting a global consumer platform.
You’ll work across cloud architecture, infrastructure as code, security engineering, reliability, automation, and disaster recovery.
A major focus will be designing and implementing cross-region disaster recovery against clearly defined RPO and RTO objectives.
You’ll also strengthen cloud security through federated identity, least-privilege access, organizational guardrails, and automated controls.
The role combines deep hands-on engineering with significant ownership over infrastructure standards, operational resilience, and technical decision-making.
You’ll work in a Terraform-first, pull-request-driven environment alongside distributed engineering and DevOps teams across multiple time zones.
This is an opportunity to take ownership of consequential infrastructure challenges while improving the resilience, security, and scalability of a sophisticated AWS estate.
Accountabilities
-
Extend and maintain a versioned Terraform module library and account customization layers covering networking, IPAM, private hosted zones, SSM access, backup infrastructure, account baselines, and other foundational services.
-
Design and maintain VPC architectures, cross-account networking, shared services, and hub-and-spoke connectivity patterns across a multi-account AWS environment.
-
Support the migration and consolidation of legacy infrastructure into standardized AWS account structures, Terraform state management, and infrastructure-as-code practices while minimizing operational risk and downtime.
-
Design and implement cross-region disaster recovery for critical platforms, including database replication, backup replication, storage replication, container image distribution, DNS failover, and infrastructure parity across regions.
-
Establish recovery objectives with business stakeholders, define RPO and RTO targets, map application dependencies into recovery tiers, and determine appropriate restoration sequences.
-
Standardize organization-wide backup policies, retention requirements, centralized backup vaults, and cross-region or cross-account backup strategies.
-
Conduct scheduled disaster recovery tests and game days, documenting repeatable runbooks that enable reliable recovery by on-call teams.
-
Design and strengthen IAM policies, trust relationships, permission boundaries, IAM Identity Center permission sets, and least-privilege access patterns.
-
Replace long-lived credentials with federated SSO, OIDC, and short-lived assumed-role access across infrastructure and CI/CD workflows.
-
Develop and maintain Service Control Policies, organizational security controls, automated remediation, and preventive guardrails.
-
Expand detective security capabilities using AWS Config, Security Hub, GuardDuty, CloudTrail, and related monitoring or SIEM tooling.
-
Manage secrets through AWS Secrets Manager and SSM Parameter Store, including secure storage and rotation processes.
-
Support security incident response and CloudTrail investigations while converting findings into durable engineering improvements.
-
Operate core infrastructure services, including monitoring, alerting, backups, patching, lifecycle management, and reliability tooling.
-
Improve Prometheus, Grafana, Alertmanager, and CloudWatch coverage so meaningful production issues are identified and escalated quickly.
-
Troubleshoot production issues across compute, networking, storage, databases, and other infrastructure layers.
-
Develop automation and operational tooling using Python and Bash, while improving infrastructure CI/CD workflows.
-
Maintain and evolve AWS Control Tower and the landing zone, while contributing to cloud cost visibility and optimization.
-
Use modern AI-assisted engineering tools, including Claude, Claude Code, and Cowork, as part of day-to-day development and infrastructure workflows.
-
5+ years of experience building and operating production AWS infrastructure.
-
Strong hands-on Terraform experience in a team environment, including module design and versioning, remote state management, drift handling, and pull-request-based workflows.
-
Proven experience with AWS Organizations or Control Tower in multi-account environments, including Service Control Policies and cross-account IAM.
-
Deep understanding of AWS IAM, including policy and trust-policy design, permission boundaries, federation, and practical least-privilege implementation.
-
Production experience with containers, particularly ECS or Kubernetes, as well as managed relational database technologies.
-
Strong Python or Bash scripting skills for automation, tooling, and operational improvements.
-
Experience migrating or consolidating legacy infrastructure while maintaining service availability and carefully sequencing technical change.
-
Demonstrated ability to balance security, reliability, cost, technical risk, and delivery speed when making infrastructure decisions.
-
Strong written communication skills, with the ability to document architecture, decisions, runbooks, incidents, and operational procedures clearly.
-
Hands-on experience using AI-assisted development and engineering tools such as Claude, Claude Code, and Cowork.
-
Experience with AWS account-vending technologies such as Account Factory for Terraform, AWS Landing Zone Accelerator, or similar solutions is preferred.
-
Experience implementing IAM Identity Center or SAML/OIDC federation with an external identity provider is preferred.
-
Experience designing and testing cross-region disaster recovery against defined RPO and RTO objectives is highly valued.
-
Experience operating Amazon Aurora at scale, including replication, failover, backup, and restoration strategies.
-
Familiarity with AWS Security Hub, GuardDuty, AWS Config, CloudTrail analysis, SIEM platforms, and/or endpoint detection and response tools.
-
Experience managing AWS Backup at organizational scale, including cross-region and cross-account backup copies.
-
Familiarity with Cloudflare or comparable CDN and WAF technologies.
-
Experience supporting data or AI workloads on AWS, such as Redshift, MWAA, or Bedrock, is a plus.
-
Experience collaborating with distributed teams across multiple time zones.
-
AWS certifications such as Solutions Architect, SysOps Administrator, or Security Specialty are preferred.
-
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
-
Availability to work three evenings per week from approximately 5:30–7:30 PM U.S. Pacific Time to overlap with distributed development and DevOps partners, with flexibility during daytime hours.
-
Willingness to travel occasionally to Las Vegas for planning sessions, incident postmortems, and team events.
-
Base salary of $130,000–$180,000, depending on experience and qualifications.
-
Remote position available anywhere in the United States.
-
Medical, dental, and vision insurance with 99% of employee premiums covered.
-
65% coverage for qualified dependents.
-
100% employer coverage for short-term disability, long-term disability, and life insurance.
-
50% 401(k) match on contributions up to 6% per month.
-
Flexible Spending Account.
-
Flexible paid time off.
-
Professional development budget supporting AWS certifications, conferences, and continued technical learning.
-
Flexibility to adjust daytime working hours around the required evening collaboration window.
-
Opportunity to work on large-scale AWS infrastructure, cloud security, automation, and cross-region disaster recovery.
-
Significant technical ownership and scope within a hands-on senior engineering role.