Roche Logo

Roche

Principal/Senior Site Reliability Engineer

Posted 7 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Design, build, and operate resilient cloud and on-prem infrastructure for MLOps and HPC at global scale. Implement Infrastructure as Code, disaster recovery, autoscaling, observability, and chaos engineering. Provide technical leadership, mentor engineers, define SLAs/SLOs/SLIs, and collaborate across teams to optimize ML and HPC workloads.
The summary above was generated by AI

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections,  where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data & Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients.

The Opportunity:

  • You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.

  • You design for resilience building disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.

  • You strengthen reliability through chaos engineering running experiments that validate systems and surface weaknesses before they become incidents.

  • You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.

  • You provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement.

  • You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.

Who you are:

  • You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.

  • You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands-on with Docker, Kubernetes, and Kubeflow.

  • You are an expert in automation scripting confidently in Python, Bash, or Go, with a strong grasp of GPU-accelerated computing and HPC workload scaling.

  • You lead through influence communicating and mentoring with clarity, and solving complex problems with a methodical approach.

  • You hold a degree in Computer Science or a related technical field or bring equivalent experience in software and site reliability engineering.

Preferred:

  • Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.

  • Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.

  • Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.

Relocation benefits are available for this position.

If building the resilient platforms that power the next generation of medicines is your calling, apply now and help accelerate science for patients worldwide.

 

 

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.


Let’s build a healthier future, together.

The statements herein are intended to describe the general nature and level of work being performed by employees, and are not to be construed as an exhaustive list of responsibilities, duties, and skills required of personnel so classified. Furthermore, they do not establish a contract for employment and are subject to change at the discretion of Roche Products Ltd. At Roche Products we believe diversity drives innovation and we are committed to building a diverse and flexible working environment. All qualified applicants will receive consideration for employment without regard to race, religion or belief, sex, gender reassignment, sexual orientation, marriage and civil partnership, pregnancy and maternity, disability or age. We recognise the importance of flexible working and will review all applicants’ requests with care. At Roche difference is valued and we are proud to be an equal opportunity employer where you are encouraged to bring your whole self to work.

Similar Jobs

9 Minutes Ago
In-Office or Remote
London, Greater London, England, GBR
Senior level
Senior level
eCommerce • Mobile • Retail
The Product Manager, International Growth will lead the growth strategy for non-US markets, focusing on user acquisition, engagement, and retention while collaborating with various teams to ensure a standardized and effective product experience.
10 Minutes Ago
Easy Apply
Remote
United Kingdom
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
As a Staff Forward Deployed Engineer at GitLab, you'll work with strategic accounts to improve platform adoption, develop reusable solutions, and influence product decisions. Responsibilities include technical discovery, architectural design, and creating reusable assets to facilitate customer outcomes.
Top Skills: Ai SystemsAnsibleCi/CdDockerGitlabGoHelmRuby On RailsTerraformYaml
10 Minutes Ago
Easy Apply
Remote
United Kingdom
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Ownership of backend features for Agentic Tools: design and implement GraphQL/REST APIs, build secure scalable Ruby on Rails services, improve RSpec automated tests, collaborate across product and AI teams, participate in Tier 2 on-call, and shape architecture for AI agent interactions with GitLab.
Top Skills: Gitlab McpGraphQLPythonRestRspecRuby On RailsVue

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account