CMC Markets Logo

CMC Markets

ML Ops Engineer

Posted Yesterday
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Mid level
In-Office
London, Greater London, England, GBR
Mid level
Build and operate ML platform capabilities that move models from experimentation into reliable production. Responsibilities include automating training, validation, deployment, retraining and rollback; operating batch and online inference services; implementing CI/CD, observability, monitoring, alerting and incident response; improving scalability, security and cost efficiency; and developing production-grade Python tooling. The role collaborates with research, software, platform, security, data and product teams.
The summary above was generated by AI

ML Ops Engineer

London

We’re hiring an ML Ops Engineer to build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services.

You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams.

This is not a research role. It is a hands-on engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response.

What you’ll work on

ML lifecycle and platform engineering

  • Build repeatable workflows for model training, validation, promotion, deployment and retraining.
  • Productionise models through packaging, versioning, model registry integration, deployment automation and safe rollback.
  • Design CI/CD pipelines for ML systems, including automated testing, validation, release controls and environment promotion.
  • Manage experiment tracking, model metadata and reproducibility across research and production.
  • Build reusable tooling and platform capabilities that support multiple models and engineering teams.

Model serving and observability

  • Deploy and operate batch and online inference services in containerised cloud environments.
  • Define and meet availability, latency, throughput and recovery objectives for ML services.
  • Monitor service health, infrastructure, data-quality signals, data drift, prediction drift and model performance decay.
  • Establish dashboards, alerting and operational runbooks so failures are detected and resolved quickly.
  • Support automated or controlled retraining, model promotion, rollback and model retirement.
  • Debug production issues across model, application, infrastructure and critical data-dependency layers.

Reliability, security and engineering quality

  • Improve system robustness, scalability and cost efficiency through automation, observability and infrastructure as code.
  • Write production-grade Python for long-running services, deployment tooling and ML workflows.
  • Establish testing, validation, release and incident-management practices for ML systems.
  • Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance.
  • Make explicit trade-offs between research flexibility, delivery speed, operational risk and production stability.

Additional responsibilities

  • Maintain personal/professional development to meet the changing demands of the role, including all relevant regulatory and legislative training
  • When dealing with all customers, clients or colleagues ensure that we provide a clear, fair and consistent high quality service that presents a professional and positive image of CMC Markets
  • Take all reasonable steps to ensure appropriate confidentiality
  • Undertake such other duties, training and/or hours of work as may be reasonably required and which are consistent with the general level of responsibility of this role

KEY SKILLS AND EXPERIENCE

  • 3–7 years’ professional experience in MLOps, ML platform engineering, ML infrastructure, backend engineering, DevOps or SRE.
  • Strong production Python skills, including clean APIs, testing, performance awareness and maintainable services.
  • Experience deploying, serving and operating machine-learning models in production environments.
  • Practical understanding of the ML lifecycle, including training, validation, inference, model release, monitoring and retraining.
  • Experience designing CI/CD workflows and release processes for ML or other production software systems.
  • Hands-on experience with at least one workflow or orchestration system used for ML training, validation or deployment.
  • Comfort working with cloud infrastructure, containers, infrastructure as code and service networking.
  • Strong understanding of observability, monitoring, alerting, incident response and common failure modes in ML systems.
  • Ability to reason about system design, reliability and operational trade-offs—not just individual tools.
  • Clear communication skills and the ability to work effectively with research, engineering, platform, security and product teams.

Nice to have

  • Prior ownership of model monitoring, drift detection or automated retraining.
  • Familiarity with model registries, feature stores and offline/online feature-consistency challenges.
  • Experience supporting multiple models, services or teams on a shared ML platform.
  • Exposure to regulated or high-reliability production environments.
  • Experience with PyTorch or similar ML frameworks and model-serving technologies.

Technology environment

Language: Python

ML tooling: PyTorch or similar frameworks, experiment tracking and model registries

Workflow orchestration: ML workflows for training, validation, deployment and retraining

Deployment: Containers, model-serving frameworks and infrastructure as code

Observability: Metrics, logging, tracing, alerting and monitoring across model, service and platform layers

Cloud: Managed compute, storage and networking, with a provider-agnostic mindset

The technology stack will evolve. We value engineers who understand why systems are designed in particular ways and can adapt as requirements and tools change.

Why this role matters

Machine-learning models only create value when they are correct, observable and dependable in production. This role is responsible for making that happen.

You’ll reduce the gap between promising experiments and production systems that can be trusted by downstream products and customers. Your work will improve the reliability, speed and scalability of the ML platform across the organisation.

If you care about operational clarity, robust engineering and building ML systems that do not silently fail, this role gives you direct leverage over the success of our machine-learning capabilities.

CMC Markets is an equal opportunities employer and positively encourages applications from suitably qualified and eligible candidates regardless of gender, sexual orientation, marital or civil partner status, gender reassignment, race, colour, nationality, ethnic or national origin, religion or belief, disability or age

HQ

CMC Markets London, England Office

London, United Kingdom

Similar Jobs

3 Hours Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Information Technology • Software • Consulting
Design and develop scalable cloud-native data platforms, batch and streaming pipelines, data lakes, and ETL/ELT solutions. Support analytics and AI/ML workloads while implementing DevSecOps, Infrastructure as Code, and reliable data services. Collaborate with clients, architects, engineers, and consultants to solve complex data challenges across defence, government, and national security environments. The role requires strong Python, SQL, Spark, AWS, containerization, and distributed data processing expertise, plus current high-level security clearance and up to 80% onsite work.
Top Skills: Amazon AthenaAmazon KinesisAmazon RedshiftAmazon S3Apache KafkaSparkAWSAws EmrAws GlueAws LambdaCi/CdDevsecopsDockerInfrastructure As CodeJavaKubernetesPythonScalaSQL
13 Days Ago
In-Office
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Transportation
Owns release gates and quality standards across Wayve’s machine learning training and delivery lifecycle. Reviews model and metric changes, validates evaluation results, identifies pipeline bottlenecks, and builds checks and automation with AI Platform and CI/CD teams. The role also improves evaluation methodology, model delivery workflows, and MLOps practices while ensuring model releases meet safety and performance standards.
Top Skills: Ci/CdGithub ActionsGrafanaMl Lifecycle ManagementMlopsModel DeploymentModel RegistryPyTorchQuantizationTensorrt
27 Days Ago
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Information Technology
Design, scale, and maintain cloud-native MLOps and LLMOps infrastructure. Provision GPU orchestration, optimise compute/network/storage, build CI/CD and MLOps pipelines, deploy LLMs with inference engines, automate with IaC, monitor model/data drift and latency, and manage cloud GPU cost and observability.
Top Skills: AnsibleArgocdAWSAzureBashDeepspeedDockerGCPGithub ActionsGoGrafanaHelmHugging Face TgiIstioJenkinsKubecostKubeflowKubernetesLangchainLangsmithLinux Kernel TuningMlflowNvidia Gpu OperatorOpentelemetryPrometheusPythonRaySlurmTensorrt-LlmTerraformTriton Inference ServerVllmWeights & Biases

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account