Senior Platform Engineer I
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Platform Engineer I based in Canada.
As a Senior Platform Engineer I, you will help build and strengthen the infrastructure that enables engineering teams to deliver stable, secure, reliable, and high-performing products. You will take hands-on ownership of Kubernetes platforms, observability, automation, and infrastructure lifecycle management. Working closely with product, security, analytics, and engineering teams, you will influence infrastructure decisions from design through deployment. You will also identify reliability and cost-efficiency opportunities before they become operational issues. The role combines deep technical execution with incident response, automation, and continuous platform improvement. This is a remote opportunity offering significant ownership and the chance to make a measurable impact across a growing technology environment.
Accountabilities
-
Own observability, logging, and alerting for Kubernetes clusters and critical workloads, ensuring platform health and reliability are clearly monitored.
-
Build and maintain automation covering the full Kubernetes cluster lifecycle, including provisioning, scaling, upgrades, and ongoing operational management.
-
Lead infrastructure cost-optimization initiatives, including workload right-sizing, scheduling improvements, and identifying and eliminating unnecessary Kubernetes spend.
-
Partner directly with engineering teams to shape infrastructure decisions throughout the architecture, development, deployment, and operational lifecycle.
-
Proactively identify reliability bottlenecks and implement durable solutions before they develop into production incidents.
-
Participate in an on-call rotation and respond quickly to service disruptions affecting users.
-
Lead incident response, root-cause investigations, and blameless postmortems for platform and infrastructure issues.
-
Contribute to the continuous improvement of internal developer and data platforms, with a focus on stability, security, performance, and developer experience.
-
Help evolve platform engineering practices through automation, infrastructure-as-code, GitOps, observability, and modern cloud-native technologies.
-
5+ years of professional experience in platform engineering, infrastructure engineering, Site Reliability Engineering, or a related discipline.
-
Demonstrated experience running Kubernetes in production at scale, including day-two operations such as troubleshooting, upgrades, scaling, and reliability management.
-
Experience working with at least one major cloud provider, such as AWS, Google Cloud Platform, or Microsoft Azure.
-
Hands-on experience with Kubernetes cluster autoscaling and workload scheduling technologies such as Karpenter or Cluster Autoscaler.
-
Proficiency with infrastructure-as-code tools such as Terraform, Crossplane, Pulumi, or equivalent technologies.
-
Practical experience with observability technologies and frameworks such as OpenTelemetry, Prometheus, and eBPF.
-
Familiarity with GitOps-based delivery and tools such as Argo CD and Kargo.
-
Experience managing and deploying Kubernetes applications using tools such as Kustomize and Helm.
-
Strong knowledge of Linux systems and low-level operating-system fundamentals.
-
Solid understanding of web and network protocols, including HTTP, TLS, and DNS.
-
Strong programming skills in at least one modern programming language beyond basic scripting, such as Go, Python, or Ruby.
-
Comfortable incorporating AI into daily engineering workflows and using AI as a collaborative development tool.
-
Strong troubleshooting, analytical, communication, and collaboration skills, with the ability to work effectively across engineering disciplines.
-
Nice-to-have experience with multi-region deployments on AWS or another cloud provider.
-
Nice-to-have knowledge of Kubernetes internals and the broader Kubernetes ecosystem, including operators, CNIs, and service meshes.
-
Nice-to-have exposure to or experience with operationalizing canary infrastructure.
-
Annual compensation of CA$156,500–CA$215,000, depending on geographic compensation zone.
-
Compensation range does not include potential bonuses, commissions, benefits, or equity that may form part of the overall compensation package.
-
Fully remote position for candidates based in Canada.
-
Flexibility to work from anywhere in Canada.
-
Unlimited paid time off.
-
Health and dental benefits.
-
Up to CA$2,025 toward home-office IT setup.
-
Up to 2% matching RRSP contributions.
-
Learning and development opportunities.
-
Up to CA$67.50 toward internet or mobile phone service.
-
Opportunity to collaborate with a team of creative technical problem-solvers and make a meaningful impact on a growing technology platform.
-
Commitment to an inclusive, safe, and harassment- and discrimination-free workplace.