Fractile Logo

Fractile

DevOps Engineer (Senior) - Bazel Monorepo

Reposted 23 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Senior level
In-Office
London, Greater London, England, GBR
Senior level
Develop cutting-edge AI accelerator solutions by improving developer workflows, managing CI/CD pipelines, optimizing Bazel builds, and extending system support.
The summary above was generated by AI
DevOps Engineer — Bazel Monorepo

Fractile  ·  London or Bristol  ·  Full-time  ·  Hybrid (3 days onsite)

About Fractile

Fractile is building silicon, systems, and software that will redefine the frontier of AI — running the world's most advanced models at radically higher speed and lower cost. Founded in 2022, our fast-growing team spans hardware and software, working hand in hand to deliver a new class of ML hardware from first principles.

If you want to build infrastructure that every engineer in the company depends on, this is the role.

The Role

We're looking for a Senior DevOps Engineer to own our Bazel monorepo and CI/CD infrastructure at Fractile — designing the tooling, pipelines, and observability that power everything we build, from ML models and kernel drivers to hardware simulation. You'll bring both deep build system expertise and an SRE mindset: owning reliability, debugging CI failures systematically, and giving the engineering organisation clear visibility into how our infrastructure is performing.

Your work will become the foundation every engineer at Fractile builds on. That's the scope, and that's the opportunity.

What You'll Work On
  • Design and own Bazel rules and extensions across the stack
  • Scale our monorepo as we grow across Python, C++, Rust, SystemVerilog, and ML workloads
  • Create reproducible, multi-language build pipelines
  • Debug and triage CI pipeline failures across a multi-language monorepo — developing systematic approaches to fault isolation, root cause analysis, and resolution in line with established SRE practices
  • Optimise CI performance across large compute clusters, applying SLO-driven thinking to build reliability and developer productivity
  • Build and maintain infrastructure monitoring and observability tooling, giving the engineering organisation clear visibility into build health, pipeline performance, and system reliability
  • Define alerting thresholds, runbooks, and incident response workflows for build and CI infrastructure
  • Define the developer experience for every engineer at Fractile
  • Contribute upstream to Bazel rules we depend on
Why This Role Stands Out
  • Extreme variety — ML, compilers, kernel drivers, simulators, hardware verification
  • High impact — your work becomes the backbone of the entire engineering organisation
  • Deep collaboration with Simulation, Runtime, and Hardware teams
  • Real ownership — you shape how Fractile builds and monitors software
What We're Looking For

You'll be a strong fit if you have:

  • 5+ years in software or infrastructure engineering
  • 3+ years working with build systems
  • Strong hands-on experience with Bazel
  • Python scripting and automation skills
  • Experience with CI/CD for large-scale products, including debugging and resolving pipeline failures at scale
  • Familiarity with infrastructure monitoring and observability tooling (e.g. Prometheus, Grafana, Datadog, or similar)
  • Working knowledge of SRE principles — SLOs, alerting design, incident management, and on-call practice
Why Fractile
  • Work at the frontier of AI hardware on genuinely novel engineering problems
  • A collaborative culture built on curiosity, initiative, and technical depth
  • Hybrid working: 3 days onsite (London or Bristol), 2 days remote
  • Modern offices in London and Bristol — choose your base
  • Competitive salary and equity in a fast-growing company

Fractile is committed to building a diverse and inclusive team. We welcome applications from people of all backgrounds and actively encourage candidates from underrepresented groups to apply.


Similar Jobs

18 Days Ago
In-Office
Square Mile, Greater London, England, GBR
Senior level
Senior level
Fintech • Analytics
The Lead DevOps Engineer will automate infrastructure, lead CI/CD pipeline design, manage Kubernetes workloads, and mentor a DevOps team while ensuring high system reliability.
Top Skills: AWSBashDatadogEksElkGitlab CiGrafanaIstioKubernetesPrometheusPythonTerraform
15 Days Ago
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Fintech
The Senior DevOps Engineer will maintain and support Opayo Web Applications, enhance automation, and ensure continuous delivery in a high-transaction environment while collaborating with various teams.
Top Skills: AnsibleAWSNew RelicPuppetScripting LanguagesSumologicTerraform
25 Days Ago
In-Office or Remote
London, Greater London, England, GBR
Mid level
Mid level
Fintech • Software • Financial Services
As a DevOps Engineer, you will ensure infrastructure reliability and security, automate deployment, administer systems, and integrate security testing tools.
Top Skills: AnsibleAWSAzureBashDatadogDockerGithub ActionsHelmJavaScriptKubernetesLinuxOpen TelemetryPrometheusPythonSastSIEMTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account