Anaplan Logo

Anaplan

ML Ops Engineer

Posted 4 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Senior level
In-Office
London, Greater London, England, GBR
Senior level
Design, scale, and maintain cloud-native MLOps and LLMOps infrastructure. Provision GPU orchestration, optimise compute/network/storage, build CI/CD and MLOps pipelines, deploy LLMs with inference engines, automate with IaC, monitor model/data drift and latency, and manage cloud GPU cost and observability.
The summary above was generated by AI

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.

What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.

Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.

Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating our wins – big and small.

Supported by operating principles of being strategy-led, values-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!

Role Overview

We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario planning platform.

You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency.

Your Impact

  • Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).
  • Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.
  • Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.
  • Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.
  • Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).
  • Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.
  • Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.
  • Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.
  • Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.
  • Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.

Your Skills

  • Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.
  • Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.
  • Demonstrated experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments.
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).
  • Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.
  • Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.

Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB)

We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. DEIB improves our workforce, enhances trust with our partners and customers, and drives business success. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper – this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We hire you for who you are, and we want you to bring your authentic self to work every day! 

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation.  

Fraud Recruitment Disclaimer  

It has come to our attention that fraudulent and fictitious job opportunities are being circulated on the Internet. Prospective candidates are being contacted by certain individuals, mainly through telephone calls, emails and correspondence, claiming they are representatives of Anaplan. The main purpose of these correspondences and announcements is to obtain privileged information from individuals.  

Anaplan does not:  

  • Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person.   
  • Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team whenever possible and then followed up via written communication.  

All emails from Anaplan would come from an @anaplan.com email address. Should you have any doubts about the authenticity of an email, letter or telephone communication purportedly from, for, or on behalf of Anaplan, please send an email to [email protected] before taking any further action in relation to the correspondence.   

Candidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.


Anaplan London, England Office

338 Euston Road, Floor 15, London, United Kingdom, NW1 3BT

Similar Jobs

18 Days Ago
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Robotics
Lead ownership of ML models and infrastructure lifecycle: develop, deploy, monitor, and improve ML systems. Build CI/CD for software and ML workflows, automate orchestration and environment management, ensure scalability, reliability, observability, and security, troubleshoot production issues, and document operational procedures while collaborating with engineers, data scientists, and product teams.
Top Skills: ArgocdAWSAzureBatch ProcessingCi/CdData LakesDockerETLGCPGithub ActionsGitlab CiHelmKubernetesPythonPyTorchStream ProcessingTensorFlowTerraform
3 Days Ago
Hybrid
London, Greater London, England, GBR
Expert/Leader
Expert/Leader
Edtech • Information Technology • Software
Lead design and roadmap for a scalable ML/GenAI platform: lifecycle architecture (training, deployment, monitoring), cloud-native GPU infrastructure, CI/CD for ML, observability and governance, LLM/vector retrieval capabilities, self-service tooling, cross-functional alignment, and mentoring senior engineers to accelerate safe, cost-efficient production ML at company scale.
Top Skills: Artifact ManagementAutoscalingAWSCi/CdDistributed ComputeFeature ManagementGCPGpuInfrastructure-As-CodeKubernetesLangchainLlamaindexModel ServingModel VersioningObservabilityPrompt EvaluationRetrieval SystemsVector Stores
3 Days Ago
In-Office
London, Greater London, England, GBR
Mid level
Mid level
Artificial Intelligence • Healthtech
Own and operate end-to-end ML infrastructure: build and maintain Airflow pipelines, CI/CD and IaC for model training/deployment, manage MLflow registry, deploy to AWS and edge devices, implement monitoring/data-versioning, ensure HIPAA/SOC2 compliance, and support reproducible, reliable ML in production.
Top Skills: Apache AirflowAws BatchAws CloudwatchAws Ec2Aws IamAws S3DockerGitInfrastructure-As-CodeMlflowPythonSnowflakeSQL

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account