Point your AI agent at freehire and let it find you a job.

Get the CLI →

Jobgether

NewBe an early applicant

Sr Platform Engineer, ML Infrastructure

Posted Updated
Discussion

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr Platform Engineer, ML Infrastructure based in United States.

This senior engineering role focuses on building the foundational platforms and infrastructure that enable machine learning teams to develop and operate systems more efficiently. You’ll design scalable tooling and services supporting the full ML lifecycle across cloud and on-premises environments. The role combines platform engineering, distributed systems, Kubernetes, and ML infrastructure in a highly technical environment. You’ll independently lead complex initiatives from architecture through production and ongoing operations. Working closely with ML and infrastructure engineers, you’ll translate technical needs into reliable, easy-to-use platform capabilities. Your work will improve developer productivity, operational excellence, scalability, and the delivery of production ML systems.

Accountabilities:

  • Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations.
  • Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently.
  • Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership.
  • Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience.
  • Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions.
  • Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency.
  • Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform.
  • Maintain high standards for software quality, production readiness, observability, and operational excellence.
  • Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases.
  • Requirements:

    • 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems.
    • Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling.
    • Proven experience designing, building, and operating production platforms used by multiple engineering teams.
    • Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations.
    • Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred.
    • Strong understanding of production reliability, observability, scalability, and operational best practices.
    • Strong technical judgment and the ability to independently drive complex initiatives from discovery through production while collaborating across technical teams.
    • Experience with developer platforms, internal tooling, or services that improve engineering productivity and reduce operational complexity is preferred.
    • Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems is a plus.
    • Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training, inference, or large-scale image processing is preferred.
    • Experience supporting production ML platforms in computer vision, robotics, or related technical domains is advantageous.
    • Demonstrated technical leadership through architecture, mentoring, or influencing technical direction across teams.
    • Benefits:

      • Base salary range of $160,000–$287,000 per year, depending on experience, qualifications, education, location, and skills.
      • Eligibility for an annual performance bonus.
      • Competitive benefits package.
      • Full-time, remote position within the United States.
      • Visa sponsorship may be available for this position.
      • Opportunities for career development, mentorship, and learning and development.
      • Inclusive and collaborative work environment focused on meaningful, technically challenging work.
      • Opportunity to work on advanced machine learning, robotics, and intelligent machinery technologies with cross-disciplinary teams.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Skills

What Senior DevOps jobs ask for — and how much of it you have →

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available