Design and build core components of the ML compiler, focusing on optimization strategies, compiler architecture, and collaboration with hardware teams.
You will design and build core components of our ML compiler, owning critical parts of the transformation and optimization pipeline.
Requirements
Benefits
Your work will directly determine whether models fit within hardware constraints and achieve real-time performance.
- Design and implement compiler architecture across:
- IR design and transformation pipelines
- Optimization strategies for scheduling, tiling, and memory reuse
- Lead development of complex compiler passes such as:
- Global memory allocation and liveness-driven reuse
- Cross-operator fusion and graph partitioning
- Hardware-aware scheduling strategies
- Develop cost models and optimization heuristics
- Explore advanced techniques:
- Constraint-based optimization (e.g., ILP/MILP/CP)
- Scheduling optimization
- Drive debugging of system-level issues (correctness, performance, HW mismatches)
- Collaborate with hardware teams on co-design of abstractions and execution models
Requirements
Required qualifications and experience:
- 4+ years of experience in compilers, systems, or performance engineering
- Masters or PhD in Computer Science, Electrical Engineering, Math, or a related field
- Deep experience with at least one:
- ML compiler frameworks (MLIR, TVM, XLA, etc.)
- Low-level optimization (scheduling, memory, tiling)
- Proven ability to design non-trivial compiler systems or passes
- Strong intuition for performance across compute, memory, and data movement
- Comfort working with hardware constraints
Nice to have:
- Experience with constraint solvers (MILP, ILP, CP)
- Background in accelerator architectures or embedded systems
- Experience optimizing ML workloads for latency/power
- Familiarity with DSP or real-time signal processing
Benefits
- 401(k)
- Medical insurance
- Vision insurance
- Dental insurance
- Commuter benefits
- Disability insurance
- Paid maternity leave
- Paid paternity leave
- Child care support
femtoAI is an equal opportunity employer committed to a diverse workforce which strives to create an inclusive working environment empowering everyone to do their best work. We do not discriminate on the basis of race, ethnicity, religion, gender, gender identity, sexual orientation, age, marital status, veteran status, or disability status.
Similar Jobs
Automotive
Design and implement compiler optimizations and new compiler features for a high-performance neural network inference platform. Analyze generated code performance, collaborate with hardware architects for hardware/software codesign, and work with model developers to tune networks for inference efficiency and accuracy. Leverage agentic AI to accelerate analysis, planning, and production software improvements.
Top Skills:
Agentic AiC++Compilers For Parallel ArchitecturesLinear Algebra ComputationMl Inference
Semiconductor
As an ML Compiler Engineer, you will develop efficient programs for ML models, implementing transformations and optimizations within the compiler stack.
Top Skills:
C++LlvmMlirOnnxPythonPyTorchTvmXla
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop algorithms and optimizations for inference and compiler stack, enhance performance metrics, collaborate with hardware teams, and publish work on novel approaches.
Top Skills:
C/C++LlvmMlirOnnxPyTorchRustTensorFlow
What you need to know about the London Tech Scene
London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.


