Member of Technical Staff, ML Inference Engineering
About the Role
Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.
We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.
What You'll Do
Inference Optimizations
● Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.
● Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.
● Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.
● Drive model-level execution improvements, including mixed precision and advanced quantization strategies.
Serving Optimizations
● Write and optimize custom inference server backends to handle complex business logic, scripting, and state management.
● Own and scale our overarching model serving infrastructure across multi-GPU, multi-node deployments to meet strict on-premise latency budgets.
● Build robust, fault-tolerant runtime services, complete with comprehensive performance benchmarking and monitoring.
Requirements
● 8+ years of experience writing high-quality, high-performance code, including at least 5 years focused on machine learning systems.
● Experience with one of C/C++/Rust and Python.
● Deep familiarity with modern NVIDIA GPU architectures (e.g., Ada Lovelace, Blackwell), CUDA, and low-level system profiling.
● Hands-on experience with ML compilers and optimization frameworks (e.g., Apache TVM, TorchInductor/Dynamo, TensorRT, Triton).
● Experience building or extending model serving infrastructure, specifically writing custom C++ backends for Triton Inference Server (or similar serving engines).
● Fluency in the AI serving stack, from kernels and quantization up to schedulers, state management, and autoscaling.
● A record of shipping research or systems that other people build on, whether in a lab or in industry.
Nice-to-have:
● Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes.
● A research-leaning or systems background in Speech (STT, TTS, S2S) or LLM inference, with work you can point to.
● Experience operating large-scale, on-premise AI training or inference clusters, including bare-metal provisioning and Kubernetes management.
● Familiarity with high-performance networking (InfiniBand or RoCE), distributed storage systems, and hardware health monitoring.
● Experience maintaining or contributing to open-source ML or systems projects.