Microsoft Logo

Microsoft

AI Reliability Engineer, HPC

Posted 3 Hours Ago
Be an Early Applicant
In-Office
London, England, GBR
Mid level
In-Office
London, England, GBR
Mid level
Build and operate reliable, highly available HPC infrastructure for AI model training and inference. Responsibilities include observability, automation, deployment, scaling, failover, incident management, postmortems, security, compliance, and collaboration with ML and platform engineering teams. The role requires operating CPU/GPU clusters, distributed systems, storage, networking, workload schedulers, and ML pipelines while optimizing capacity and costs.
The summary above was generated by AI
Overview

As Microsoft continues to push the boundaries of AI, we are on the lookout for passionate individuals to work with us on the most interesting and challenging AI questions of our time. Our vision is bold and broad — to build systems that have true artificial intelligence across agents, applications, services, and infrastructure. It’s also inclusive: we aim to make AI accessible to all — consumers, businesses, developers — so that everyone can realize its benefits.

We’re looking for an experienced AI Reliability Engineer to join our High Performance Computing (HPC) infrastructure team. In this role, you’ll blend software engineering and systems engineering to keep our large-scale distributed AI infrastructure reliable and efficient. You’ll ensure that AI systems stay efficient and reliable with very high uptimes.


Microsoft AI   Our mission is to build AI that amplifies human potential and empowers people around the world. We strive to deliver breakthroughs that advance science, education, productivity, and global well-being.   We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!   MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.


Responsibilities
  • Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
  • Observability: Design and maintain monitoring, alerting, and logging systems to provide real-time visibility into all aspects of HPC systems including GPU, clusters, storage and networking.
  • Automation & Tooling: Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments.
  • Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements.
  • Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments.
  • Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.
  • Embody our Culture and Values.

Qualifications

Required Qualifications:

  • Bachelor’s Degree in Computer Science, or related technical discipline AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering OR equivalent experience

Preferred Qualifications:

  • Master’s Degree in Computer Science, or related technical discipline AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
  • OR equivalent experience
  • Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads.
  • Experience with public cloud platforms such as Azure, AWS, or GCP, including infrastructure-as-code.
  • Experience with monitoring and observability tools such as Grafana, Datadog, or OpenTelemetry.
  • Programming or scripting experience in Python, Go, or Bash.
  • Experience with distributed systems, networking, storage, and high-performance computing (HPC).
  • Experience operating GPU clusters and workload schedulers for ML/AI workloads.
  • Experience with ML training or inference pipelines.
  • Experience with capacity planning and cost optimization for GPU-based infrastructure.

Software Engineering IC5 - The typical base pay range for this role across United Kingdom is £ 93,500.00 - £ 161,800.00 per year. Certain roles may be eligible for benefits and other compensation.

Find additional benefits and pay information here:
https://careers.microsoft.com/v2/global/en/corporate-pay/united-kingdom-corporate-pay.html


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Microsoft London, England Office

2 Kingdom St, London, United Kingdom, W2 6BD

Similar Jobs

16 Minutes Ago
Hybrid
London, Greater London, England, GBR
Expert/Leader
Expert/Leader
Financial Services
Leads development and operation of secure, high-performance Java trading systems. Designs latency-sensitive distributed services, APIs, integration strategies, resiliency, observability, and automation. Oversees architecture, coding, testing, incident response, production stability, and SDLC security. Guides AI-assisted engineering practices, mentors developers, and coordinates technical planning and delivery across globally distributed teams.
Top Skills: APIsAutomated TestingCi/CdDistributed SystemsFixJava 17+KafkaKubernetesLinuxMicroservicesMqObservabilityPythonResilience EngineeringSolaceSpringSpring Boot
17 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Financial Services
Leads the strategy, roadmap, development, and performance of risk tooling products supporting control assurance, audit readiness, risk reporting, and operational risk management. Drives automation of manual risk and controls processes, partners with Engineering, Risk, Audit, Controls, and Architecture teams, establishes governance for complex initiatives, and advances enterprise analytics and AI-enabled risk solutions. Coaches product teams, manages stakeholders, monitors market trends, and ensures product investments deliver measurable business outcomes.
Top Skills: AIAnalytics DashboardsData MeshData PlatformsEnterprise ReportingMachine LearningWorkflow Automation
17 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Financial Services
Serves as a strategic partner and chief of staff to the Product Management leader, driving engineering roadmaps, OKR alignment, portfolio governance, resource and budget management, executive reporting, and organizational processes. Uses AI-assisted planning and reporting, prepares data-driven recommendations, partners with HR and Finance, supports talent management, and coordinates cross-functional delivery across senior stakeholders.
Top Skills: Enterprise-Authorized AiExcelMicrosoft Powerpoint

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account