Diffractive Labs Logo

Diffractive Labs

Data Engineer

Posted 28 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Mid level
In-Office
London, Greater London, England, GBR
Mid level
Design and own data architecture and ingestion pipelines that convert wet-lab instrument outputs into versioned, reproducible datasets for ML and RL training. Build tooling for dataset inspection, lineage, and integration with model training workflows, and collaborate closely with ML researchers to support pretraining, midtraining, and RL experiments.
The summary above was generated by AI

What We're Looking For

We are seeking a Data Engineer to build and own the infrastructure that underpins our AI-driven materials discovery platform. You'll work directly with world-renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. As a Data Engineer, you will own the data architecture that powers this research.

This role requires a builder who understands both large-scale machine learning pipelines and the messy reality of physical lab data. You will serve as the critical link between experimental results generated at the bench and the models evaluating them, ensuring our research team always has the exact, high-quality datasets required to push the frontier of materials science. We are looking for someone with a rigorous, experimental mindset who thrives in an interdisciplinary environment and operates with a high degree of technical ownership.

What You'll Do

  • Drive the overarching data architecture across our training stack, mapping out data requirements with ML researchers and evaluating new external sources to fill knowledge gaps.

  • Design and deploy the ingestion pipelines that capture physical experimental data directly from our wet lab instruments and feed it seamlessly into our model training workflows.

  • Construct robust, reproducible systems for processing, standardizing, and versioning diverse scientific corpora, creating a highly reliable foundation for the research team.

  • Build internal tooling that allows machine learning researchers and physical scientists to effectively query, inspect, and audit the data feeding into pretraining, midtraining, and RL runs.

  • Continuously integrate emerging techniques in synthetic data generation, data selection, and data-efficient training into our production systems.

Skills & Qualifications

  • 3+ years of engineering experience focused on large-scale data pipelines, ideally within an applied ML, scientific, or LLM training environment.

  • High proficiency in Python and modern workflow orchestration frameworks (e.g., Dagster, Airflow, Prefect, or similar).

  • Demonstrated experience with dataset lineage, versioning, and reproducibility tooling (such as DVC, Delta Lake, or custom equivalents).

  • A track record of collaborating directly with machine learning researchers, translating complex modeling needs into scalable pipeline architecture and back again.

  • Strong DevOps fundamentals, including hands-on experience with containerization (Docker, Kubernetes) and CI/CD deployment.

Nice to Have

  • Prior experience processing and structuring data from physical laboratory instrumentation, computational simulations, or multimodal scientific sources.

  • A background in curating datasets for domain-specific continued pretraining or instruction tuning.

  • An academic or practical background in physics, materials science, or chemistry.

Why Join Us

You are building the foundation for breakthroughs in magnetic materials that will directly influence the future of energy and computing hardware.

Diffractive is building the AI Material Scientist that autonomously learns from real-world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter.

You'll join a small, high-calibre team where your work has real impact from day one. We're London-based with a flexible approach to how and where you work. We offer competitive salary, generous equity and benefits. You'll have a real stake in what you build and in the company's overall success.

How to Apply

If you're excited about this role and believe you could thrive in it, we'd encourage you to apply even if you may not align with every part of the job description.

Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities.

If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.

Hit the apply button below to submit your application. We are looking forward to hearing from you!

Similar Jobs

16 Days Ago
Easy Apply
Hybrid
London, England, GBR
Easy Apply
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Design, build, deploy, and optimize distributed big-data ingestion, standardization, and ML pipelines at scale. Support platform reliability, cost controls, security accreditations, and customer integrations. Serve as technical lead in client meetings, mentor engineers, and recommend tools and best practices for data pipeline development and deployment.
Top Skills: AWSAzureDatabricksDatadogDockerEc2GCPGitIamKubernetesOpensearchPostgresPythonRest ApisS3SparkSQLTerraform
22 Days Ago
In-Office
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain near-real-time and batch regulatory reporting systems. Migrate legacy SQL Server solutions to GCP BigQuery, design cloud-native data models and ETL/ELT pipelines, implement validations, monitoring, CI/CD and IaC, apply AI-assisted tooling, perform QA and documentation, and collaborate with senior engineers to ensure accurate, traceable regulator-ready outputs.
Top Skills: Claude CodeGCPGitGitlabGoogle BigqueryGoogle Cloud ComposerGoogle Cloud FunctionsGoogle Cloud StorageGoogle Pub/SubMicrosoft Sql ServerSQLTerraform
22 Days Ago
In-Office
Mid level
Mid level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain near-real-time and batch regulatory reporting systems. Migrate legacy SQL Server reporting to GCP BigQuery, design cloud-native data models and ETL/ELT pipelines, implement validation, monitoring, CI/CD and IaC, and apply AI-assisted tooling to improve accuracy, traceability and performance for regulatory submissions.
Top Skills: Ci/CdClaude CodeCloud ComposerCloud FunctionsCloud StorageEltETLGCPGitGitlabGoogle BigqueryMicrosoft Sql ServerPub/SubSQLTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account