Point your AI agent at freehire and let it find you a job.

Get the CLI →

Gimlet Labs

Member of Technical Staff - Compiler Engineer

Posted Updated
Discussion

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

About the role

As a Member of Technical Staff, Compiler, you will build Gimlet's AI Compiler Stack: the layer that determines how AI workloads are represented, optimized, and executed across hardware with fundamentally different architectures, memory systems, and performance characteristics.

The compiler is one of the hardest and most critical parts of our inference stack. Its decisions reach directly into scheduling, communication, memory movement, kernel execution, and end-to-end serving performance of Gimlet’s inference stack. A compiler built for a single target can bake in assumptions about the memory hierarchy and execution model. We can't. That constraint is what makes the work interesting.

Our stack is built on open-source MLIR and LLVM compiler infrastructure. We design our own dialects, passes, and lowering paths on top of that infrastructure rather than maintaining a proprietary compiler from scratch. You will be working in an ecosystem you can carry with you, alongside people who know it deeply.

What you will work on:

  • Scaling the compiler architecture to a growing set of heterogeneous compute architectures. New CPUs, GPUs, and accelerators arrive with different memory models, execution models, and toolchains, and each one can't cost us a rewrite. You will work on the abstractions, dialects, and lowering paths that let the stack absorb a new target quickly and still generate code competitive with a vendor-native toolchain. Our architecture for running AI workloads across heterogeneous hardware describes where the compiler sits in the broader system.

  • Optimizing end-to-end performance that includes implementing operator fusion, target-agnostic middle-end optimizations, target-specific optimizations and the execution-planning decisions that shape how computation is partitioned across devices and how intermediate state moves between them. This also includes evaluating and integrating agent-driven kernel generation as a first-class code generation strategy: deciding when a generated kernel beats a library or hand-written one, and how the compiler makes that call automatically and safely. We have published early results on AI-generated CUDA kernels and AI-generated Metal kernels; folding this into the compiler itself is open ground.

  • Supporting new models and serving techniques in production. New architectures and inference strategies should run efficiently without per-target hand-tuning. Our work on low-latency speculative decoding and prefill/decode disaggregation across accelerators shows the shape of these problems end to end.

Who you will work with

This is a small team where the boundaries between disciplines are thin. You will be part of the compiler team and will work directly with engineers and researchers working on LLM architecture, inference optimization, kernel engineering, distributed systems, and datacenter infrastructure deciding what gets served and on what hardware. Compiler decisions here are made with full visibility into the workload and the machine, which is rarer than it sounds.

What success looks like

In your first 12–18 months, you will:

  • Ship compiler optimizations that measurably improve production latency, throughput, utilization, or cost.

  • Help bring a new accelerator architecture into production.

  • Reduce the engineering work required to support additional hardware targets.

  • Design execution strategies for partitioning workloads across heterogeneous hardware.

  • Extend how Gimlet selects between generated, library, and hand-written kernels.

  • Enable a new model architecture or serving technique across multiple targets.

You may be a good fit if you have

  • Experience building ML compilers and runtimes with sharp focus on performance optimizations.

  • Experience with MLIR, LLVM, or comparable compiler infrastructure.

  • Experience designing intermediate representations, writing compiler passes, or implementing lowering and code-generation pipelines.

  • Strong C++ and/or Python skills.

  • A strong understanding of SSA, memory systems, scheduling, and hardware efficiency.

  • Experience with operator fusion, tiling, layout transformations, or memory planning.

  • Experience using profilers to investigate performance across software and hardware boundaries.

  • Experience with IREE, XLA, TVM, Triton, or similar ML compilers and programming tools.

Strong candidates may also have

  • Experience optimizing ML inference or model-serving workloads.

  • Familiarity with kernel dispatch, launch APIs, memory allocators, or runtime interfaces.

  • Experience across GPU and non-GPU accelerator architectures.

  • Experience with automated kernel generation, autotuning, or search-based optimization.

  • Contributions to open-source compiler infrastructure.

Why join now?

Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.

  • Solve hard problems.

  • Own meaningful work.

  • Build for production.

  • Help define what’s next.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Skills

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available