Point your AI agent at freehire and let it find you a job.

Get the CLI →

Sanas

Member of Technical Staff, ML Inference Engineering

Posted Updated
Discussion

About the Role

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Inference Optimizations

Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.

Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.

Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.

Drive model-level execution improvements, including mixed precision and advanced quantization strategies.

Serving Optimizations

Write and optimize custom inference server backends to handle complex business logic, scripting, and state management.

Own and scale our overarching model serving infrastructure across multi-GPU, multi-node deployments to meet strict on-premise latency budgets.

Build robust, fault-tolerant runtime services, complete with comprehensive performance benchmarking and monitoring.


Requirements

8+ years of experience writing high-quality, high-performance code, including at least 5 years focused on machine learning systems.

Experience with one of C/C++/Rust and Python.

Deep familiarity with modern NVIDIA GPU architectures (e.g., Ada Lovelace, Blackwell), CUDA, and low-level system profiling.

Hands-on experience with ML compilers and optimization frameworks (e.g., Apache TVM, TorchInductor/Dynamo, TensorRT, Triton).

Experience building or extending model serving infrastructure, specifically writing custom C++ backends for Triton Inference Server (or similar serving engines).

Fluency in the AI serving stack, from kernels and quantization up to schedulers, state management, and autoscaling.

A record of shipping research or systems that other people build on, whether in a lab or in industry.

Nice-to-have:

Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes.

A research-leaning or systems background in Speech (STT, TTS, S2S) or LLM inference, with work you can point to.

Experience operating large-scale, on-premise AI training or inference clusters, including bare-metal provisioning and Kubernetes management.

Familiarity with high-performance networking (InfiniBand or RoCE), distributed storage systems, and hardware health monitoring.

Experience maintaining or contributing to open-source ML or systems projects.

Skills

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available