Designs and develops highly available cloud services for AI inference workloads. Responsibilities include building Kubernetes controllers and platform capabilities for deployment, scheduling, scaling, recovery, upgrades, observability, and lifecycle management. The role leads production readiness reviews, resolves complex infrastructure issues, improves reliability and performance, drives architecture and design reviews, mentors engineers, and influences technical direction across teams.
As an engineer on Arm's AI Inference Cloud team, you will shape the technical direction and develop highly available, scalable services for running AI inference workloads. You will guide architecture and actively contribute to development across Kubernetes orchestration, workload management, service delivery, and observability. Partnering with AI compute, Inference Runtime, and product teams to enhance the performance and usability of Arm's AI platform.
Responsibilities:
Necessary Skills and Experience:
Preferred Skills and Experience:
In Return:
You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities for AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. Your work will directly impact the success of our AI projects, shape Arm's AI inference capabilities, defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!
(Recruiter to Complete)
Additional Information:
Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Responsibilities:
- Define and build the architecture for cloud-based AI inference services.
- Develop Kubernetes controllers and platform capabilities supporting workload deployment, scheduling, recovery, scaling, upgrades, and lifecycle management.
- Establish production practices for health validation, progressive rollout, rollback, observability, and service objectives. Improve platform reliability, scalability, performance, and resource efficiency.
- Lead production readiness reviews and resolve complex issues across services, Kubernetes, networking, and compute infrastructure. Turn incidents and operational bottlenecks into durable platform improvements.
- Lead design and build reviews, mentor engineers, and drive technical alignment across teams.
Necessary Skills and Experience:
- 5+ years of experience, or equivalent proven impact, building distributed systems, cloud platforms, or production infrastructure.
- Deep production experience with Kubernetes, including controllers, operators, scheduling, resource management, networking, and workload lifecycle management.
- Strong software and production engineering skills, including hands-on programming in Go, C++, Rust, Python, or a similar language, and experience with reliable services, APIs, concurrency, observability, deployment safety, capacity planning, and incident response.
- A track record of leading complex technical initiatives while remaining hands-on, including architecture, implementation, debugging, mentoring, and influencing technical direction across teams.
- Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.
Preferred Skills and Experience:
- Experience with AI infrastructure, model serving, or accelerator-backed workloads.
- Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.
- Knowledge of inference performance and resource-efficiency considerations.
In Return:
You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities for AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. Your work will directly impact the success of our AI projects, shape Arm's AI inference capabilities, defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!
(Recruiter to Complete)
Additional Information:
Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.
Salary Range:
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at Arm
At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [email protected] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.
Hybrid Working at Arm
Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.
Equal Opportunities at Arm
Arm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don't discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Similar Jobs at Arm
Artificial Intelligence • Internet of Things • Semiconductor
Build and operate secure, scalable AI compute platform services, including backend APIs, asynchronous workflows, Kubernetes controllers, and multi-tenant workload capabilities. The role focuses on distributed training and inference, platform reliability, observability, developer tooling, and infrastructure simplification. Responsibilities include collaborating across engineering teams, improving scalability and developer productivity, and providing technical leadership for production AI platform capabilities.
Top Skills:
APIsAsynchronous ProcessingContainersGitopsGoGrafanaKafkaKubernetesKubernetes OperatorsOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
Artificial Intelligence • Internet of Things • Semiconductor
Design and build secure, scalable platform services for distributed AI training and inference. Develop backend APIs, Kubernetes controllers, orchestration systems, asynchronous workflows, multi-tenant access controls, and developer tools. Improve reliability, observability, scalability, and engineering productivity while partnering with infrastructure and AI teams. The role also provides technical leadership, mentoring, and cross-functional collaboration throughout platform development and production delivery.
Top Skills:
ContainersGitopsGoGrafanaKafkaKubernetesOpentelemetryPostgresPrometheusPythonPyTorchRayRelational DatabasesRpcVllm
Artificial Intelligence • Internet of Things • Semiconductor
Design and lead development of robotics software foundations connecting sensing, planning, control, actuation, middleware, operating systems, and hardware accelerators. Build APIs, SDKs, diagnostics, testing workflows, and production-ready reference platforms for Arm-powered devices. Optimize real-time performance, reliability, safety, and resource usage while mentoring engineers and collaborating across robotics, AI, hardware, product, and customer teams.
Top Skills:
APIsC++CpusDspsGpusIpcLinuxNpusPythonReal-Time SystemsRos 2SdksSecure Boot
What you need to know about the London Tech Scene
London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

