Microsoft Logo

Microsoft

Compute Orchestration & Scheduling

Posted Yesterday
Be an Early Applicant
In-Office
Enfield, Middlesex, England, GBR
Senior level
In-Office
Enfield, Middlesex, England, GBR
Senior level
Build and scale compute infrastructure for AI workloads across large GPU clusters. Responsibilities include cluster orchestration, workload scheduling, resource allocation, quota management, topology-aware placement, runtime initialization, fault recovery, observability, and utilization optimization. The role partners with researchers and model-development teams to deliver reliable infrastructure for distributed training and inference.
The summary above was generated by AI
Overview

Microsoft AI is looking for engineers to build the compute infrastructure powering frontier-model development. The Compute Orchestration & Scheduling team owns cluster orchestration, workload scheduling, resource allocation, quota management, and the systems that enable fast initialization and reliable fault recovery on next-generation GPU supercomputers. 

Our team prepares infrastructure for the next generation of distributed training and inference across multiple locations and a rapidly growing fleet of accelerators, including NVIDIA Grace Blackwell, Vera Rubin, and AMD GPU platforms, alongside the CPU, network, and storage systems they depend on. 

Key challenges include topology-aware placement, large-scale distributed training, fast and reliable initialization of training runtimes, rapid recovery for long-running AI jobs, observability into cluster behavior, and improved utilization across increasingly large and diverse accelerator fleet. Better scheduling efficiency and platform reliability translate directly into more effective compute for AI research and products. 

You will work closely with researchers, model engineers, hardware architects, and infrastructure teams to turn frontier-model requirements into scalable platform capabilities. We value engineers who navigate ambiguity, remove roadblocks, and deliver improvements to users quickly and iteratively.


Responsibilities
  • Develop and tune the compute infrastructure stack allocating NVIDIA Grace Blackwell (GB) and Vera Rubin (VR) resources to AI workloads. 

  • Scale GPU clusters across hardware generations to thousands of accelerators and beyond. 

  • Use operational data and workload insights to inform the compute (GPU and CPU) roadmap for large-scale AI research. 

  • Partner with model-development teams to improve the infrastructure used to train and serve AI models. 

  • Find practical ways around roadblocks and deliver improvements rapidly, iterating with users in a fast-paced, design-driven environment. 

  • Embody Microsoft’s culture and values. 


Qualifications

Required Qualifications

  • A bachelor’s degree in computer science or a related technical field and at least six years of engineering experience writing code in languages such as C, C++, Python, Go, or JavaScript; or equivalent practical experience. 

Preferred Qualifications

  • Significant additional engineering experience, with a master’s degree or equivalent practical experience, building production software and distributed systems.
  • Experience with Ray, Kubernetes, Kueue, Volcano, or another AI-focused system for scheduling, scaling, or fault tolerance is especially relevant. 


Software Engineering IC5 - The typical base pay range for this role across United Kingdom is £ 93,500.00 - £ 161,800.00 per year. Certain roles may be eligible for benefits and other compensation.

Find additional benefits and pay information here:
https://careers.microsoft.com/v2/global/en/corporate-pay/united-kingdom-corporate-pay.html


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Microsoft London, England Office

2 Kingdom St, London, United Kingdom, W2 6BD

Similar Jobs

2 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Define and deliver AI-powered analytics and decision-support products. Set strategy and roadmaps, partner with engineering and analytics teams to build LLM/RAG-driven conversational and agentic experiences, translate customer needs into requirements, use data to measure performance, and ensure trust, governance, privacy, and responsible AI practices.
Top Skills: Ai AgentsConversational AiGenerative AiLarge Language Models (Llms)Machine LearningOrchestration FrameworksRetrieval-Augmented Generation (Rag)
3 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Manage globally distributed crisis operations analysts across 24/7 regional pods. Responsibilities include staffing and capacity planning, SLA and response-quality oversight, crisis escalation, war-room coordination, coaching, onboarding, performance management, AI adoption, operational data analysis, SOP and runbook maintenance, post-event debriefs, reporting, and process improvement. The role requires calm judgment during natural hazards, geopolitical crises, aviation disruptions, public health events, and other emerging threats.
Top Skills: Ai-Assisted MonitoringConfluenceCrisis24DtnEverbridgeExcelJIRALarge Language ModelsMicrosoft 365Microsoft TeamsNc4OsacOutlookQuerybookServicenowSharepointStormgeoThe Weather CompanyTrinoWorkflow Automation
3 Minutes Ago
Hybrid
London, Greater London, England, GBR
Mid level
Mid level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Leads continuous-improvement initiatives for air traveler servicing, using operational data, traveler feedback, and frontline insights to identify problems and design practical improvements. Partners with Product, Technology, Operations, Content, Training, Analytics, and airline stakeholders to coordinate requirements, testing, launches, and adoption. Establishes success metrics, monitors service outcomes, manages risks and dependencies, and communicates recommendations and progress to stakeholders and leaders.

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account