Anaplan Logo

Anaplan

ML Ops Engineer

Posted 27 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Senior level
In-Office
London, Greater London, England, GBR
Senior level
Design, scale, and maintain cloud-native MLOps and LLMOps infrastructure. Provision GPU orchestration, optimise compute/network/storage, build CI/CD and MLOps pipelines, deploy LLMs with inference engines, automate with IaC, monitor model/data drift and latency, and manage cloud GPU cost and observability.
The summary above was generated by AI

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.

What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.

Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.

Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating our wins – big and small.

Supported by operating principles of being strategy-led, values-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!

Role Overview

We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario planning platform.

You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency.

Your Impact

  • Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).
  • Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.
  • Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.
  • Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.
  • Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).
  • Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.
  • Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.
  • Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.
  • Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.
  • Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.

Your Skills

  • Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.
  • Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.
  • Demonstrated experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments.
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).
  • Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.
  • Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.

Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB)

We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. DEIB improves our workforce, enhances trust with our partners and customers, and drives business success. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper – this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We hire you for who you are, and we want you to bring your authentic self to work every day! 

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation.  

Fraud Recruitment Disclaimer  

It has come to our attention that fraudulent and fictitious job opportunities are being circulated on the Internet. Prospective candidates are being contacted by certain individuals, mainly through telephone calls, emails and correspondence, claiming they are representatives of Anaplan. The main purpose of these correspondences and announcements is to obtain privileged information from individuals.  

Anaplan does not:  

  • Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person.   
  • Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team whenever possible and then followed up via written communication.  

All emails from Anaplan would come from an @anaplan.com email address. Should you have any doubts about the authenticity of an email, letter or telephone communication purportedly from, for, or on behalf of Anaplan, please send an email to [email protected] before taking any further action in relation to the correspondence.   

Candidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.


Anaplan London, England Office

338 Euston Road, Floor 15, London, United Kingdom, NW1 3BT

Similar Jobs

4 Hours Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Information Technology • Software • Consulting
Design and develop scalable cloud-native data platforms, batch and streaming pipelines, data lakes, and ETL/ELT solutions. Support analytics and AI/ML workloads while implementing DevSecOps, Infrastructure as Code, and reliable data services. Collaborate with clients, architects, engineers, and consultants to solve complex data challenges across defence, government, and national security environments. The role requires strong Python, SQL, Spark, AWS, containerization, and distributed data processing expertise, plus current high-level security clearance and up to 80% onsite work.
Top Skills: Amazon AthenaAmazon KinesisAmazon RedshiftAmazon S3Apache KafkaSparkAWSAws EmrAws GlueAws LambdaCi/CdDevsecopsDockerInfrastructure As CodeJavaKubernetesPythonScalaSQL
Yesterday
In-Office
London, Greater London, England, GBR
Mid level
Mid level
Fintech • Payments • Financial Services
Build and operate ML platform capabilities that move models from experimentation into reliable production. Responsibilities include automating training, validation, deployment, retraining and rollback; operating batch and online inference services; implementing CI/CD, observability, monitoring, alerting and incident response; improving scalability, security and cost efficiency; and developing production-grade Python tooling. The role collaborates with research, software, platform, security, data and product teams.
Top Skills: AlertingCi/CdCloud ComputingContainersExperiment TrackingFeature StoresInfrastructure As CodeLoggingMetricsModel RegistriesModel-Serving FrameworksMonitoringPythonPyTorchTracing
13 Days Ago
In-Office
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Transportation
Owns release gates and quality standards across Wayve’s machine learning training and delivery lifecycle. Reviews model and metric changes, validates evaluation results, identifies pipeline bottlenecks, and builds checks and automation with AI Platform and CI/CD teams. The role also improves evaluation methodology, model delivery workflows, and MLOps practices while ensuring model releases meet safety and performance standards.
Top Skills: Ci/CdGithub ActionsGrafanaMl Lifecycle ManagementMlopsModel DeploymentModel RegistryPyTorchQuantizationTensorrt

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account