Fractile Logo

Fractile

Developer Operations / DevOps — Bazel Monorepo

Reposted 17 Days Ago
Be an Early Applicant
In-Office
Bristol, England
Senior level
In-Office
Bristol, England
Senior level
The role involves improving developer workflows, managing CI/CD pipelines, optimizing Bazel builds, and extending support for new languages.
The summary above was generated by AI

Developer Operations / DevOps — Bazel Monorepo
Location: Bristol / London

About Fractile

Fractile was founded in 2022 on the bet that, eventually, the world's most capable AI systems would be limited in their impact by the time taken to produce useful outputs. We bet everything on the logical conclusion: that the only way to truly unlock this latent value, to make speed viable at scale, was to radically re-invent the hardware that we run our frontier AI models on. Ever since, we have been building chips and systems that tackle this problem: how to efficiently generate output at thousands of tokens per second, while handling the complexity and capacity challenges of operating large models at very long contexts.

The workloads that push to the limits of the current frontier are already transformational; it is the technical and economic limits on inference speed that are constraining progress. The defining work of the 21st century will be marked by the engine of inference delivering immense and diffuse chains of intellectual inquiry, in drug discovery, in software engineering, in materials discovery, in any field where progress is driven by deep reasoning and intelligence to resolve complex problems.

The Role

About the Software organisation at Fractile

Developer Experience sits within the Software organisation at Fractile, which is responsible for developing a full software stack for our groundbreaking AI inference systems. That's everything from ML compilers, device drivers and systems firmware, application level runtime and ecosystem integrations, ML and compute libraries, great developer tooling and a full portfolio of simulators, through to datacenter scale workload deployment solutions. At Fractile, we know that a fantastic software stack is a critical and central part of any AI inference solution and it sits at the heart of everything we're doing.

About the team and role

The DevOps team at Fractile owns the Bazel monorepo and CI/CD infrastructure that every engineer in the company depends on — designing the tooling, pipelines, and observability that power everything we build, from ML models and kernel drivers to hardware simulation. The team brings both deep build system expertise and an SRE mindset: owning reliability, debugging CI failures systematically, and giving the engineering organisation clear visibility into how our infrastructure is performing.

As a DevOps Engineer you will design and own Bazel rules and extensions across the stack, scale our monorepo as we grow across Python, C++, Rust, SystemVerilog, and ML workloads, and create reproducible multi-language build pipelines. You will debug and triage CI pipeline failures, optimise CI performance across large compute clusters, build and maintain infrastructure monitoring and observability tooling, and define alerting thresholds, runbooks, and incident response workflows for build and CI infrastructure. You will contribute upstream to Bazel rules we depend on and help shape the developer experience for every engineer at Fractile.

About you

You bring deep build system expertise alongside an SRE mindset. You are systematic in how you approach fault isolation and root cause analysis, and you apply SLO-driven thinking to build reliability and developer productivity. You thrive on high variety — ML, compilers, kernel drivers, simulators, hardware verification — and you understand that your work becomes the foundation every engineer at Fractile builds on. You take that responsibility seriously.

Key Requirements

  • 5+ years in software or infrastructure engineering
  • 3+ years working with build systems
  • Strong hands-on experience with Bazel
  • Python scripting and automation skills
  • Experience with CI/CD for large-scale products, including debugging and resolving pipeline failures at scale
  • Familiarity with infrastructure monitoring and observability tooling (e.g. Prometheus, Grafana, Datadog, or similar)
  • Working knowledge of SRE principles — SLOs, alerting design, incident management, and on-call practice

What We Offer

  • Competitive salary: A competitive salary reflective of your experience and the specialist nature of the role
  • Equity & Ownership: meaningful equity so everyone shares in the value creation
  • Benefits: Private Medical, Dental and Vision, Contributory Pension, 25 Days holiday plus bank holidays and Life/Critical Illness Insurance
  • Diverse & fun office: we believe the hardest problems get solved by the broadest range of minds. We are committed to Equal Employment Opportunity through attracting and retaining a diverse team and building an inclusive environment

Fractile is seeking to increase the clock speed of global progress, one chip at a time. We've recently raised $220M from investors including Founders Fund and Accel and our most important work lies ahead. Join us!

Export controls

Our work involves technologies subject to UK, US and other international export control regulations. Certain roles may require additional eligibility checks to ensure compliance with applicable law. We'll be transparent about this throughout the hiring process.

Similar Jobs

17 Days Ago
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Semiconductor
Develop cutting-edge AI accelerator solutions by improving developer workflows, managing CI/CD pipelines, optimizing Bazel builds, and extending system support.
Top Skills: BazelCC++CocotbDockerGithub ActionsIcarus VerilogKubernetesPythonRustSystemverilogVerilator
A Minute Ago
Easy Apply
Hybrid
London, Greater London, England, GBR
Easy Apply
Senior level
Senior level
Big Data • Cloud • Software • Database
As a Senior Solutions Architect at MongoDB, you'll design scalable systems, guide customers in leveraging MongoDB technology, and collaborate with sales teams to drive solutions while ensuring customer success.
Top Skills: C#C++JavaMongoDBNode.jsPythonSQL
28 Minutes Ago
Remote or Hybrid
United Kingdom
Entry level
Entry level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Analyzes adversary intrusions, malware, files, and events to create and improve machine-learning security detections. Reviews false positives and negatives, investigates emerging threats, analyzes data at scale, and coordinates with internal teams on product and service improvements. The role also addresses customer questions about detection-model efficacy and may require on-call coverage several times annually.
Top Skills: AssemblyCC++JavaMachine LearningPublic Cloud InfrastructureWindows ApiWindows Os

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account