Allwyn UK Logo

Allwyn UK

Senior / Lead Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
In-Office
Watford, Hertfordshire, England, GBR
Senior level
In-Office
Watford, Hertfordshire, England, GBR
Senior level
Lead SRE responsible for reliability across customer-facing systems using SLOs/SLIs/error budgets. Own incident command and on-call, drive automation (Terraform, CI/CD), transition from ECS to EKS, optimise capacity and performance for high-concurrency events, implement observability (Splunk, CloudWatch, Grafana, Quantum Metric), prioritise SRE backlog, and mentor teams while reporting reliability metrics to leadership.
The summary above was generated by AI

At the heart of everything we do is our vision to change lives every day, and our mission to grow The National Lottery responsibly and champion its impact.  


We are Allwyn UK, part of the Allwyn Entertainment Group – a multi-national lottery operator with a market-leading presence across the USA (Michigan and Illinois) and Europe, including Czech Republic, Austria, Greece, Cyprus and Italy. 


While the main contribution of The National Lottery to society is through the funds to good causes, at Allwyn we put our purpose and values at the heart of everything we do.  Join us as we embark on a once-in-a-lifetime, largescale transformation journey by creating a National Lottery that delivers more money to good causes.   


We’ll talk a bit more about us further down the page, but for now – let’s talk about the role and who we’re looking for… 


A bit about the role 

At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events.

The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams.

Objectives of the role

  • Own reliability outcomes across services using SLOs, SLIs, and error budgets
  • Improve availability, latency, and scalability across Instant-Win and Draw-based platforms
  • Lead incident response and operational readiness, including peak jackpot events
  • Drive automation and platform maturity, reducing manual operational effort
  • Establish clear reporting on reliability, incidents, and service health trends
  • Supporting with thought-leadership and developing long-term roadmap

What you’ll be doing 

Reliability engineering & technical leadership

  • Define and govern SLOs / SLIs / error budgets across critical services
  • Lead reliability design reviews across:
    • Web & mobile platforms
    • Player Identity & Protection systems
    • CMS and Geolocation services
  • Drive architecture improvements for resilience (failover, degradation, scaling patterns)

Production operations & incident leadership

  • Act as incident commander for major incidents and high-severity events
  • Lead 1-in-4 on-call rotation, covering:
    • Out-of-hours incident diagnosis
    • Peak jackpot proactive monitoring
  • Own end-to-end incident lifecycle:
    • Detection → triage → resolution → post-incident review
  • Ensure blameless post-mortems with clear remediation ownership

Observability & service insight

  • Define and evolve observability strategy using:
    • Splunk (log analytics)
    • CloudWatch (AWS telemetry)
    • Grafana (metrics visualisation)
    • Quantum Metric (user behaviour insight)
  • Standardise:
    • Alerting quality and signal-to-noise ratio
    • Dashboards aligned to SLOs and customer impact
    • Drive correlation between technical signals and user experience

Automation & platform engineering

  • Lead automation of operational processes using Terraform and scripting
  • Improve deployment and release safety (CI/CD, progressive delivery patterns)
    • Key contributor to transition strategy for ECS → EKS (Kubernetes adoption)
  • Reduce operational toil through tooling, self-healing, observability and platform improvements
  • Empower Level-1 operational teams with safe, controlled access to the tools they need to operate autonomously

Capacity & performance engineering

  • Own capacity planning for:
    • High-concurrency draw events
    • Traffic spikes during jackpots
  • Lead performance optimisation:
    • Latency reduction
    • Throughput scaling
    • Cost efficiency (AWS utilisation and associated log costs, observability license consumption)

Backlog ownership & reporting

  • Own and prioritise the SRE backlog, balancing:
    • Reliability improvements
    • Technical debt
    • Automation opportunities to reduce/offload toil
  • Produce structured reporting covering:
    • SLO performance
    • Incident trends and MTTR
    • Platform health and risk areas
  • Provide clear updates to engineering leadership and business stakeholders

Collaboration & culture

  • Embed SRE practices across engineering teams
  • Mentor engineers and SREs on:
    • Reliability engineering
    • Observability
    • Incident management
  • Promote a culture of automation, measurement, and continuous improvement
  • Supporting with thought-leadership and developing long-term SRE roadmap for Digital Operations

What experience we’re looking for 

 Technical

  • Strong experience in cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes)
  • Hands-on experience with:
    • Terraform (Infrastructure as Code)
    • Observability stacks (Splunk, CloudWatch, Grafana)
  • Strong programming skills (Python, Go, or similar)

SRE practices

  • Proven experience implementing:
    • SLOs, SLIs, error budgets
    • Incident management frameworks
    • Observability strategies
  • Strong experience in distributed systems troubleshooting

Operational

  • Experience in on-call production environments
  • Demonstrated leadership during high-severity incidents

Desirable Experience

  • Experience migrating container platforms (ECS → EKS/Kubernetes)
  • Experience supporting high-scale consumer platforms
  • Familiarity with real-time analytics / customer experience tooling (e.g., Quantum Metric)
  • Experience in regulated or high-availability environments

About us 

At Allwyn, we are dedicated to changing lives and growing the National Lottery responsibly, championing its positive impact on people, places, and the planet. 


  • Innovation - We pride ourselves on it! We’re constantly looking for new ways to excite our customers, bringing new products to market to enjoy which is all supported by our responsible play values and making them accessible to all.  
  • Giving back – Did you know that playing the lottery generates around £30m a week for charities and good causes in the UK? Our aim is to have doubled this number by the end of the first 10-year license. 
  • Sustainability – Our aim is to become a net zero national lottery. We have 2030 targets to decarbonise our operations and energy. We’ve already transitioned to renewable energy providers, made our London and Watford offices zero gas, and ensured our fleet consists of low-emission vehicles. In addition, we’re working with our value chain partners to develop a net zero target date. 
  • Empowering every voice – We believe in creating a culture where everyone feels they belong, can be themselves, has access to opportunities and can thrive for the benefit of good causes.  Our diverse teams are working hard to make all parts of The National Lottery inclusive – whether people play a game in a store or online, because when everyone can play, everyone wins.. 

An inclusive reward offering with wellbeing at the centre 

At Allwyn, inclusion is built into how we care for our people. Our benefits and policies support colleagues and their families at every stage of life and career. By prioritising wellbeing and belonging, we create a workplace where everyone feels valued, rewarded, and empowered to succeed. Our people are more than colleagues - they’re winners, driving positive change and making a real difference in communities. 


Benefits 

  • Company Bonus Scheme 
  • Matched pension contributions up to 8.5% 
  • 26 days annual leave + 2 Life Days (and bank holidays) 
  • Single Private Health Cover 
  • Complimentary Private Medical 
  • Income Protection  
  • Flexible Benefits – EV Scheme, Money Coach, Will Writing, Mortgage Advice, Dental and Eye Care Schemes. 
  • Enhanced Family Leave (Maternity, Paternity, Adoption) 
  • Wellness Allowance £500 
  • Employee Assistance Programme 
  • Discounted Health Assessments 
  • Volunteering Days 
  • Matched Funding 

We are a Disability Confident Leader which means we’ve taken proactive steps to ensure our workplace is accessible and inclusive for disabled and neurodivergent colleagues and candidates. As part of this we offer an interview to disabled applicants who meet the essential requirements of the job. 


If you need any assistance or adjustments to this job description or in the application process, please contact a member of the talent team at [email protected] and we’ll be happy to help. 


HQ

Allwyn UK Watford, England Office

Tolpits Ln, Watford, United Kingdom, WD18 9RN

Allwyn UK London, England Office

One Connaught Place, 5th Floor, London, United Kingdom, W2 2ET

Similar Jobs

43 Minutes Ago
Hybrid
London, England, GBR
Mid level
Mid level
Fintech • Mobile • Payments • Software • Financial Services
Build and own the backend infrastructure that processes all sending transactions. Collaborate with product, design, and data teams to deliver customer-facing send features, drive architectural decisions, and ensure reliability and performance of core send components.
45 Minutes Ago
Hybrid
London, Greater London, England, GBR
Junior
Junior
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Sell Square solutions to UK SMBs through full sales cycle from prospecting to close. Build and manage pipeline, exceed targets, use consultative value-selling, collaborate with cross-functional teams for onboarding, and leverage Salesforce and business development tactics (cold calling) to drive deals.
Top Skills: Salesforce
45 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Fintech • Mobile • Payments • Software • Financial Services
Lead product analytics for Wise Platform's Cash Management, building models, KPIs, dashboards and decision frameworks to drive liquidity, settlement and FX products. Deliver analyses, opportunity sizing, scenario modelling, and scalable data products to inform product roadmap and increase partner revenue and margins.
Top Skills: AirflowDbtGitLightdashNumpyPandasPlotlyPythonScipySnowflakeSQL

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account