Diffractive Labs Logo

Diffractive Labs

Platform Engineer

Posted One Month Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Mid level
Hybrid
London, Greater London, England, GBR
Mid level
Build and own cloud infrastructure and GPU compute environments for ML model training and deployment. Implement IaC, CI/CD, orchestration, containerisation, monitoring, security, and internal tooling to enable reproducible, scalable research-to-production ML workflows.
The summary above was generated by AI

What We're Looking For

We are seeking a Platform Engineer to build and own the infrastructure that underpins our AI-driven materials discovery platform. You'll work directly with world-renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible.

This is a foundational hire. You'll set the patterns others build on.

You will be joining a small, highly ambitious team of world-renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that.

What You'll Do

  • Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure-as-code; Terraform, Pulumi, or equivalent.

  • Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation.

  • Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments.

  • Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability.

  • Ensure reproducibility across research and production environments through containerisation and rigorous environment management.

  • Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking.

  • Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning.

  • Build internal tooling and documentation that lets researchers self-serve infrastructure without waiting on you.

Skills & Qualifications

  • 4+ years in a DevOps, Platform Engineering, or SRE role.

  • Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM).

  • Hands-on experience with infrastructure-as-code and container orchestration (Kubernetes or equivalent).

  • Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar.

  • Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack.

  • Deep Linux systems knowledge and strong networking fundamentals.

  • A bias for building things properly the first time, even under early-stage constraints.

Nice to Have

  • Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training).

  • Familiarity with MLOps tooling:

  • Experiment tracking (MLflow, Weights & Biases).

  • Workflow orchestration (Airflow, Prefect, Argo).

  • Data versioning (DVC).

  • Background in scientific computing or HPC environments.

  • Prior experience at a deep tech or computational science company.

Why Join Us

  • Work directly on infrastructure that enables AI to make real scientific discoveries.

  • Shape how we build from day one, no legacy systems, no inherited mess.

  • Collaborate with world-class researchers across materials science and machine learning.

Diffractive is building the AI Material Scientist that autonomously learns from real-world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter.
You'll join a small, high-calibre team where your work has real impact from day one. We're London-based with a flexible approach to how and where you work. We offer competitive salary, generous equity and benefits. You'll have a real stake in what you build and in the company's overall success.

How to Apply

If you're excited about this role and believe you could thrive in it, we'd encourage you to apply even if you may not align with every part of the job description.

Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities.

If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.

Hit the apply button below to submit your application. We are looking forward to hearing from you!

Similar Jobs

6 Days Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain release, compliance, and source-code-management platforms that enable safe, traceable software deployments. Develop automation and internal applications using Python, Go, React, and TypeScript; manage Terraform infrastructure and containerized services on Kubernetes and GCP; create dashboards and alerts; improve platform availability; support on-call escalation; and use AI tools to reduce repetitive work.
Top Skills: AlloydbArtifactoryBashDockerGCPGitlabGkeGoGrafanaKubernetesLinuxLlm PlatformsPostgresPrometheusPrompt EngineeringPythonReactSourcegraphSplunkTerraformTypescriptWindows
6 Days Ago
Hybrid
Mid level
Mid level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain release, compliance, and source-code-management platforms using Python, Go, React, and TypeScript. Automate release approvals, software-change auditing, token rotation, and repetitive engineering tasks. Manage Terraform infrastructure and Dockerized services on Kubernetes and GCP, monitor platform reliability with dashboards and alerts, resolve incidents during on-call rotations, and improve application availability. The role also involves AI-native engineering, security and compliance, and mentoring colleagues.
Top Skills: AlloydbArtifactoryBashDockerGCPGitGitlabGkeGoGrafanaKubernetesLinuxLlm PlatformsPostgresPrometheusPrompt EngineeringPythonReactSourcegraphSplunkTerraformTypescriptWindows
6 Hours Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Fintech • Mobile • Payments • Software • Financial Services
Design, build, test, and maintain highly scalable backend services and REST APIs for Wise’s pay-in platform. Develop payment orchestration, card pay-in services, PostgreSQL data systems, and Kafka-based event-driven architectures on AWS. Improve performance, reliability, security, and compliance for financial systems, participate in code reviews and on-call rotations, and collaborate cross-functionally on product delivery.
Top Skills: Apache KafkaAWSJavaKotlinPostgresRest ApisSpring BootSQL

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account