Member of Technical Staff | ML Systems
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | ML Systems based in Brazil.
This role sits at the intersection of machine learning, distributed systems, data infrastructure, and production engineering.
You’ll join an ML Systems team responsible for turning research outputs into reliable, reproducible, and governed model releases.
Your work will support researchers and platform engineers by creating the infrastructure and tooling they need to experiment and ship efficiently.
You’ll work on GPU performance, distributed training, data materialization, experiment tracking, model lineage, and release governance.
The role combines deep technical ownership with a strong internal-product mindset, where platform users and measurable engineering outcomes guide priorities.
You’ll help ensure models can run reliably across cloud, production, batch, and customer environments while remaining fully auditable.
It’s an opportunity for a strong systems engineer to have broad technical scope and directly shape the foundations of an ML platform.
Accountabilities:
-
Build high-performance CUDA kernels and compute primitives supporting the training and serving of graph neural networks.
-
Evolve distributed sampling and training infrastructure, including neighbor sampling and performance improvements for multi-node workloads.
-
Develop and maintain the systems used to efficiently train and serve machine learning models at scale.
-
Define binary and columnar data formats, including Lance, Arrow, and CSR/CSC representations, and own data materialization and feature backfills for training and evaluation.
-
Establish data contracts and consumption requirements in collaboration with teams responsible for customer and proprietary datasets.
-
Build and operate experiment tracking, checkpointing, and evaluation infrastructure with reproducibility as a default requirement.
-
Own model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models.
-
Define and operate release gates that ensure every production, batch, or on-premise model corresponds to an authorized and governed release.
-
Make model releases fully auditable by maintaining traceability across the data, code, configuration, and supporting evidence used to produce them.
-
Improve time-to-experiment by making data, compute, and experiment tracking rapidly accessible to research teams.
-
Improve time-to-governed-release by creating efficient and reliable paths from validated candidate models to production-ready releases.
-
Optimize training throughput and GPU utilization across large-scale foundation model workloads.
-
Build internal platform capabilities with the mindset of a product, focusing on the needs and productivity of researchers and platform engineers.
-
Maintain high standards for reliability, reproducibility, performance, and operational quality across ML systems.
-
Strong systems engineering background combined with production-quality Python development skills.
-
Professional experience with distributed machine learning training, including technologies such as Ray or PyTorch Distributed.
-
Experience operating or developing multi-node GPU workloads and understanding the challenges of distributed compute.
-
Experience working with columnar data formats and large-scale data materialization pipelines.
-
Familiarity with ML lifecycle infrastructure, including experiment tracking, model registries, evaluation systems, checkpoints, and reproducibility tooling.
-
Strong understanding of software engineering principles for building reliable, maintainable, production-grade infrastructure.
-
Ability to think of internal platforms as products with real users, requirements, feedback loops, and measurable outcomes.
-
Ability to take ownership of systems and outcomes rather than focusing narrowly on individual implementation tasks.
-
Strong analytical and problem-solving skills, with the ability to work across data, compute, model, and infrastructure layers.
-
A product-oriented mindset and interest in enabling researchers and engineers to work more effectively.
-
Data science expertise is not required, provided you have strong systems engineering capabilities and an understanding of ML infrastructure.
-
Experience developing CUDA kernels or optimizing GPU performance is a strong advantage.
-
Knowledge of graph neural networks, graph sampling, or large-scale graph workloads is a plus.
-
Experience with Lance, Arrow, or comparable columnar or indexed storage technologies is beneficial.
-
Familiarity with multi-cloud GPU infrastructure, including tools such as SkyPilot, is advantageous.
-
Experience with model governance, auditability, or regulated production environments—particularly financial services—is a plus.
-
Fully remote position based in Brazil.
-
Full-time opportunity within an engineering organization focused on high-impact ML infrastructure.
-
Broad technical ownership across ML systems, distributed computing, GPU performance, data infrastructure, and model governance.
-
Opportunity to work on foundational machine learning systems supporting research and production workloads.
-
Direct impact on experiment velocity, training performance, model reproducibility, and governed releases.
-
Opportunity to work with advanced technologies including CUDA, distributed GPU workloads, graph neural networks, columnar data systems, and ML lifecycle infrastructure.
-
Environment that values technical depth while giving engineers ownership of systems and outcomes.
-
Internal-platform mindset, with researchers and platform engineers treated as real users whose productivity drives priorities.
-
Opportunity to shape reliable ML infrastructure designed to operate across cloud, production, batch, and customer environments.