OrgVue Jobs

Senior Site Reliability Engineer

OrgVue

Senior Site Reliability Engineer

Reposted 17 Days Ago

Be an Early Applicant

In-Office

London, Greater London, England, GBR

Senior level

In-Office

London, Greater London, England, GBR

Senior level

The Principal Site Reliability Engineer will lead SRE transformations, ensuring system reliability and scalability, mentor engineers, and drive infrastructure as code practices.

The summary above was generated by AI

Orgvue is a leading organizational design and planning software platform that captures the power of data visualization and modelling to build more adaptable, and better performing organizations. HR, finance and business leaders use Orgvue for actionable insight and analysis that helps them make faster workforce decisions in a constantly changing world.

Orgvue is used by the world’s largest and best-known enterprises and management consulting firms to visualize and confidently build the businesses they want tomorrow, today. The company is headquartered in London, with offices in Philadelphia, The Hague, Toronto, and Sydney.

We are seeking a Principal Site Reliability Engineer who will be a senior technical leader focused on scaling and hardening our AWS- and Kubernetes-based infrastructure.

Role

In this role you will work across product, platform, and operations teams to ensure our systems are reliable, observable, and resilient, even at scale.

This role combines hands-on technical capability with strategic vision, helping us build a world-class reliability culture and a robust engineering foundation for growth. We're looking for someone who has technical expertise, is a great communicator and enjoys collaborating across multiple teams.

Responsibilities

Define and enforce SLOs, SLIs, and error budgets across critical services
Crafting and implementing a cloud infrastructure and tooling strategy
Work across our Org to level up SRE practices
Help implement robust observability metrics, logs & traces using our observability tool
Guide the team in building automated, self-healing systems
Own and evolve our incident response processes, including on-call practices and post-mortem culture
Mentor engineers across the org on best practices in reliability, operational readiness, and scalable infrastructure
Drive Infrastructure as Code (IaC) using Terraform, Kubernetes, CloudFormation and GitOps practices
Collaborate closely with security, DevOps, and software teams to ensure compliance, scalability, and operational excellence
Evaluate and introduce tools, patterns, and practices that improve the performance and reliability of our SaaS platform

Requirements

Demonstrable experience leading SRE transformations
Deep hands-on expertise with Kubernetes (EKS preferred) in production environments
Strong experience with AWS core services (EC2, EKS, RDS, S3, ALB/NLB, IAM, CloudWatch, etc.)
Expert in Infrastructure as Code using tools such as Terraform, with knowledge of GitOps workflows
Strong background in observability: metrics, visualization, logging, and tracing
Understanding of automation, SDLC, CI/CD pipelines, deployment automation, and blue/green or canary releases
Proven experience with incident management, disaster recovery planning, root cause analysis, and post-incident reviews

Benefits

Hybrid working - 1+ days a week in the London office

Wellbeing: Sanctus Coaching, Virtual fitness sessions, Wellbeing webinars, Annual Wellbeing day

Subsidised Gym Membership
Private Medical Insurance (including Dental and Vision) and Life Assurance
25 days holiday (increasing to 30 days at a rate of 1 extra day per year)
Employer pension contribution of 5% of your gross salary, if you contribute a minimum of 3%
Season ticket Loan
Cycle to Work Scheme
Annual Discretionary Bonus

'Here at Orgvue we promote individualism and a diverse workforce to build on our future success'

Similar Jobs

iManage

Senior Site Reliability Engineer

3 Days Ago

Hybrid

London, Greater London, England, GBR

Senior level

Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software

The Senior Site Reliability Engineer will automate processes, collaborate across teams, and enhance service resilience in a cloud-native environment, focusing on system scalability and best practices.

Top Skills: AksAzureBashChefDockerEfkElkGoGrafanaJavaKubernetesPowershellPrometheusPythonRubyTerraform

Experian

Senior Site Reliability Engineer

5 Days Ago

Hybrid

Senior level

Big Data • Marketing Tech • Analytics

The Senior Site Reliability Engineer will enhance system reliability, manage AWS infrastructure, automate processes, and respond to incidents while collaborating with teams to improve overall system performance.

Top Skills: AWSBashCloudFormationCloudwatchDynatraceGrafanaPrometheusPythonSplunkTerraform

Mindera

Senior Site Reliability Engineer

6 Days Ago

In-Office

London, Greater London, England, GBR

Senior level

Mobile • Software

The role involves designing and standardizing cloud platforms for AI-driven products, managing infrastructure as code, ensuring reliability, and implementing best practices across multiple projects.

Top Skills: AWSCi/CdEcsEksKongTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.