ClearRoute Logo

ClearRoute

AI Platform Engineer

Posted 22 Days Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Mid level
Hybrid
London, Greater London, England, GBR
Mid level
Design, build, and operate internal developer platforms and production-ready infrastructure. Manage Kubernetes clusters, IaC (Terraform/Pulumi/CDK), CI/CD pipelines, secrets management (Vault), and SRE/security practices. Support AI/ML platform needs (GPU scheduling, model serving, MLOps, vector DBs) and lead client-facing technical discovery, documentation, and coaching.
The summary above was generated by AI

About Us

ClearRoute is an engineering consultancy bridging Quality Engineering, Cloud Platforms and Developer Experience. We help enterprises reliably bring high-impact digital products to market faster, cheaper, and safer, working with technology leaders facing complex business challenges.

We take as much pride in our people, culture and work-life balance as we do in making better software. We’re not just making better software. We’re making the making of software better. Collaborative, entrepreneurial and dedicated to problem solving, we bring the step change our customer need to sustain innovation. Our values challenge us to do the best we can for ClearRoute, our customers and most importantly our team. This is an opportunity for you to build the organisation from the ground up, use your voice to drive change and help transform organisations and problem domains.

The Role

We are looking for a Platform Engineer to join client-facing delivery teams and help design, build, and operate modern developer platforms. You will work across a range of industries and tech stacks, so adaptability matters as much as expertise.

On any given engagement you might be building a goldenpath CI/CD pipeline, hardening a Kubernetes cluster, migrating secrets management to Vault, or running a platform engineering workshop with a client's engineering teams. You will be expected to lead technical workstreams, pair with client engineers, and leave behind well-documented, production-ready infrastructure.

Increasingly, our clients are asking us to help them build the foundations for AI, from model-serving infrastructure and MLOps pipelines to safely integrating LLM powered tooling into existing developer workflows. You don't need to be an ML engineer, but you should be curious about this space and comfortable building the platform layer that makes AI workloads production ready.

What You'll Do

Platform & Infrastructure

• Design and deliver internal developer platforms (IDPs) that improve developer experience and accelerate software delivery.

• Build and maintain infrastructure-as-code using Terraform, Pulumi, or CDK and enforce code review and testing standards.

• Manage and optimise Kubernetes clusters (EKS, GKE, AKS) including multi-tenancy, networking, RBAC, and cost controls.

• Own CI/CD pipelines end-to-end: from source control policies through build, test, security scanning, artefact management, and deployment.

• Implement secrets management and certificate lifecycle automation using HashiCorp Vault or equivalent.

Reliability & Security

• Embed SRE practices: SLOs, error budgets, runbooks, on-call design, and blameless post-mortems.

• Integrate security tooling (SAST, DAST, dependency scanning, policy-as-code) into delivery pipelines.

• Design and test disaster-recovery strategies; automate them where possible.

• Ensure compliance with client security standards and relevant regulatory frameworks.

AI & Emerging Technology

• Design and operate infrastructure for AI/ML workloads: GPU node pools, model-serving runtimes (Triton, vLLM, BentoML), and vector database deployments (pgvector, Weaviate, Qdrant).

• Build and maintain MLOps pipelines model training, versioning, evaluation, and promotion to production using platforms such as Kubeflow, MLflow, or cloud-native equivalents.

• Integrate LLM APIs and AI agent frameworks into existing developer platforms, including prompt management, observability, cost controls, and rate-limit guardrails.

• Advise clients on AI readiness: data infrastructure, governance, security controls (model access policies, output filtering), and the organisational changes that sit alongside the technical work.

• Stay current with the fast-moving AI tooling landscape and bring relevant ideas back to the team and to clients.

Client Engagement

• Lead technical discovery sessions and platform assessments with client engineering and architecture teams.

• Translate client requirements into clear technical plans and communicate trade-offs to both technical and non-technical stakeholders.

• Coach and upskill client platform and application engineers through pairing, workshops, and code review.

• Produce high-quality documentation, architecture decision records (ADRs), and runbooks that clients can own after the engagement.

What We're Looking For....

Essential

• Solid hands-on experience with at least one major cloud provider (AWS preferred; GCP or Azure accepted).

• Production experience with Kubernetes and familiarity with the surrounding ecosystem (Helm, ArgoCD / Flux, Karpenter, Cilium, etc.).

• Strong infrastructure-as-code skills, Terraform is the baseline; other tools are a bonus.

• Practical experience designing or operating CI/CD systems (GitHub Actions, GitLab CI, Tekton, or similar).

• Comfort working in Linux environments and writing automation in Python, Bash, or Go.

• The ability to explain complex technical concepts clearly to a mixed audience.

• A consulting mindset: you care about solving the client's actual problem, not just delivering a deliverable.

Ideal (not required)

• Hands-on experience with AI/ML infrastructure: GPU scheduling on Kubernetes, model-serving runtimes, or MLOps tooling (Kubeflow, MLflow, Ray, etc.).

• Familiarity with LLM APIs (OpenAI, Anthropic, Bedrock, Vertex AI) and patterns for building reliable, observable AI-powered applications.

• Experience with vector databases or semantic-search infrastructure.

• Experience with event-streaming platforms such as Apache Kafka.

• Familiarity with configuration management tools (Chef, Ansible, or similar).

• Knowledge of mainframe environments or hybrid-cloud patterns.

• Relevant certifications: CKA/CKAD, AWS Solutions Architect, HashiCorp Vault, etc.

• Prior consultancy or client-facing delivery experience.

At ClearRoute, we believe diverse perspectives lead to better outcomes, and inclusion creates the conditions for everyone to thrive. We are proud to have built a family friendly working environment and have many employees who have caring responsibilities alongside work. We welcome applications from people who require flexibility and will be happy to discuss needs on an individual basis.

We are committed to fostering a culture where all team members feel respected, supported, and empowered to do their best work. We celebrate individuality and our differences and understand that some differences may mean that you require changes made to the interview process. We are happy to cater to your needs to make the interview accessible, if this is something you require please let us know by emailing us at [email protected]

ClearRoute London, England Office

1 Waterhouse Sq, London, United Kingdom, EC1N 2ST

Similar Jobs

10 Days Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Financial Services
Design, build, and operate reliable, scalable AI/ML platform infrastructure. Develop observability, resilience, security, and automation tooling, own NFRs, participate in on-call rotations, troubleshoot production issues, and collaborate cross-functionally to support large-scale AI deployments.
Top Skills: Ai-Assisted Development ToolsAWSAzureCi/CdDockerDynatraceGCPGrafanaKubernetesOpentelemetryPythonTerraform
Yesterday
In-Office
London, Greater London, England, GBR
Entry level
Entry level
Information Technology • Consulting
Design, automate, secure, and operate AWS and Azure cloud infrastructure supporting AI and generative AI platforms. Build Terraform-based infrastructure, manage networking, identity, containers, Kubernetes, monitoring, governance, and cost controls. Support production operations, incident response, troubleshooting, documentation, and compliance in partnership with AI, Data, Security, and Application teams.
Top Skills: Amazon BedrockAmazon EksAnsibleAWSAzure Ai FoundryAzure AksBashCi/CdCopilot StudioGithub CopilotKubernetesLinuxAzureMicrosoft CopilotOpentelemetryPowershellPythonTerraformWindows
Yesterday
In-Office
London, Greater London, England, GBR
Expert/Leader
Expert/Leader
Information Technology • Software
Lead technical direction for an AI platform connecting workloads to multiple LLM providers. Build API gateway and routing infrastructure, agent orchestration, capability lifecycle tooling, access controls, API key management, caching, cost and latency optimization, adoption analytics, and trust, safety, and governance controls. Define API standards and promotion workflows while supporting auditable, scalable AI services across client engagements.
Top Skills: AnthropicApi GatewaysApi KeysApigeeAws BedrockEnvoyGoKongLlmsMcpOidcOpenaiPythonRbac

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account