SunCore Digital Logo

SunCore Digital

Senior Site Reliability Engineer (Disaster Recovery

Sorry, this job was removed at 01:01 a.m. (GMT) on Friday, Aug 07, 2026
Be an Early Applicant
Hybrid
London, England, GBR
Senior level
Hybrid
London, England, GBR
Senior level

Similar Jobs

11 Minutes Ago
In-Office
London, Greater London, England, GBR
Internship
Internship
Information Technology • Software • Financial Services • Quantitative Trading
As a PhD intern, you'll develop trading strategies, conduct statistical analysis, backtest models, and collaborate with experienced team members in quantitative research.
Top Skills: C++PythonR
11 Minutes Ago
In-Office
London, Greater London, England, GBR
Mid level
Mid level
Information Technology • Software • Financial Services • Quantitative Trading
Build and optimize low-latency, high-performance C++ trading systems for crypto markets; partner with researchers and traders to productionize strategies; optimize latency, throughput, reliability; implement pricing, risk, and execution logic; analyze large-scale market and trade data; ensure resilience, testing, and scalability across global venues.
Top Skills: C++C++17Cloud DeploymentsConcurrencyLinuxNetworking
12 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Cloud • Software
Manage and grow an Endpoint engineering team to build and scale backend infrastructure that ingests, aggregates, and stores agent-collected telemetry. Collaborate with product and cross-functional teams to deliver assurance use cases, drive data processing and availability at scale, and support engineers' career development.
Top Skills: Agile (Scrum)AIBackend InfrastructureCloudData ProcessingDistributed SystemsNetwork TelemetryTest-Driven Development
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
The summary above was generated by AI

Senior Site Reliability Engineer (Disaster Recovery

About the Role

SunCore Digital is seeking a hands-on Senior Site Reliability Engineer (Disaster Recoveryr to assess, implement, test, and document service and data recovery capabilities.

This is an engineering position, not solely a policy, compliance, or coordination role. The successful candidate will identify recovery risks, build and harden shared recovery capabilities, automate recovery-specific processes where practical, and prove that critical systems and data can be restored. Once system-specific procedures are tested and documented, the owning teams will be trained to execute and maintain them.

The engineer may draw on application developers when recovery improvements require changes to application code, databases, data flows, or system-specific behavior. Recovery procedures must ultimately be understood and maintainable by the teams that own the affected systems.

Responsibilities

Recovery assessment

  • Inventory critical applications, services, databases, storage systems, queues, infrastructure, and external recovery dependencies.

  • Determine what is currently backed up and how those backups are created, retained, protected, and monitored.

  • Identify systems with missing, incomplete, unverified, or person-dependent recovery processes.

  • Assess risks related to data loss, service loss, infrastructure failure, configuration loss, credential availability, and third-party dependencies.

  • Distinguish configured backups from proven recoverability.

  • Identify manual recovery procedures and undocumented knowledge.

  • Document confirmed capabilities, untested capabilities, and unknowns.

Backup and restoration

  • Design and implement backup improvements for critical data and configuration, then transfer ongoing operation to the designated long-term owner.

  • Translate approved recovery, retention, security, and access requirements into backup controls and monitoring.

  • Create, test, and harden restoration procedures with system owners, then transfer routine execution and system-specific maintenance to those teams.

  • Execute restoration tests using representative environments and data.

  • Validate backup completeness and integrity.

  • Measure recovery duration and potential data loss during exercises.

  • Establish monitoring and escalation so failed or incomplete backups are detected by the accountable long-term owner.

  • Automate backup verification, restoration validation, recovery evidence collection, and other repeatable recovery processes where practical; transfer routine operation of system-specific automation to the owning teams after it is tested and documented.

A backup will not be treated as reliable solely because a scheduled job reports success. Restoration must be tested.

Service recovery and failover

  • Design and implement shared service-recovery and failover capabilities, transfer supporting shared components to the designated long-term owner, and ensure system-owning teams retain responsibility for application-specific recovery behavior and procedures after handoff.

  • Identify the infrastructure, application, database, network, access, and vendor dependencies required during recovery.

  • Establish safe recovery sequencing.

  • Develop procedures for partial outages, regional failures, infrastructure loss, data corruption, and service dependency failures.

  • Implement shared recovery mitigations directly and coordinate application-specific changes with system owners, who retain responsibility for those changes.

  • Test restored services for technical functionality.

  • Work with QA to validate that critical business workflows and data remain correct after recovery.

  • Document conditions under which recovery, rollback, or failover may be unsafe.

Recovery objectives and planning

  • Work with engineering and business leadership to document recovery requirements for critical systems.

  • Translate approved business requirements into technical recovery capabilities.

  • Measure actual recovery performance against defined expectations.

  • Identify where current architecture cannot meet required recovery expectations.

  • Provide technical options and evidence to support leadership decisions.

  • Maintain a recovery dependency map for critical systems.

Final recovery priorities, risk acceptance, and business requirements remain with leadership.

Disaster-recovery exercises

  • Plan and execute recovery exercises.

  • Develop test scenarios for service, infrastructure, data, credential, and dependency failures.

  • Ensure exercises produce objective evidence.

  • Record results, defects, recovery times, data-loss observations, and unresolved risks.

  • Implement shared recovery corrective actions and coordinate system-specific corrective work with the owning teams.

  • Repeat exercises after material changes.

  • Ensure procedures can be followed by qualified staff who did not author them.

Documentation and knowledge transfer

  • Establish recovery runbook standards and create initial system-specific runbooks with the owning teams.

  • Document required access, tools, credentials, dependencies, procedures, validation steps, and escalation paths.

  • Clearly label untested procedures.

  • Train system owners and relevant operational staff to execute and maintain the system-specific procedures they own.

  • Reduce reliance on undocumented individual knowledge.

  • Ensure application teams understand their ongoing recovery responsibilities.

  • Maintain shared recovery evidence and exercise results; system owners maintain the accuracy of their service-specific runbooks.

Security and collaboration

  • Work with Security Operations to review recovery implementations that affect production access, sensitive data, credentials, networks, infrastructure, or security controls.

  • Work with SRE on service dependencies, failure behavior, observability, and operational response.

  • Define recovery requirements for infrastructure-as-code, environment recreation, artifacts, and delivery tooling; work with DevOps and platform staff to implement those requirements within their shared capabilities.

  • Work with application engineers on application-specific recovery changes.

  • Work with QA on post-recovery functional and data validation.

  • Escalate recovery risks that cannot be mitigated within current architecture or resources.

Initial Priorities

  • Inventory existing backup and recovery capabilities.

  • Identify critical systems with no confirmed recovery path.

  • Verify who owns each backup and recovery process.

  • Determine when critical backups were last restored successfully.

  • Establish repeatable restore testing.

  • Document initial recovery requirements and dependencies.

  • Implement the highest-priority recovery mitigations.

  • Create and test recovery runbooks.

  • Identify person-dependent and manual recovery processes.

  • Produce a factual recovery-readiness assessment before broader external use.

Required Qualifications

  • Five or more years of experience in disaster recovery, infrastructure resilience, backup and recovery engineering, SRE, cloud infrastructure, systems engineering, or a related technical field.

  • Hands-on experience implementing backup, restore, recovery, and failover capabilities.

  • Experience performing recovery exercises rather than only writing recovery plans.

  • Experience recovering databases, cloud infrastructure, applications, configurations, and storage systems.

  • Experience with recovery automation and scripting.

  • Understanding of recovery-time and recovery-point concepts.

  • Experience documenting and validating recovery dependencies.

  • Familiarity with cloud security, identity, secrets, networking, and access requirements during recovery.

  • Experience troubleshooting complex system failures.

  • Ability to work with application developers on recovery-related code and data changes.

  • Strong runbook and technical-documentation skills.

  • Ability to identify unknowns and avoid treating untested procedures as proven.

  • Comfort working in a small organization where recovery practices are still being established.

Bonus Points

  • Experience preparing a platform for its first external users.

  • Experience establishing a disaster-recovery capability from an early stage.

  • Experience with data-intensive, financial, digital-asset, telemetry, or operational systems.

  • Experience with infrastructure as code.

  • Experience with controlled disaster-recovery or resilience exercises.

  • Experience recovering event-driven or distributed systems.

  • Experience in a remote, asynchronous environment.

What Success Looks Like

  • Critical systems and data have documented recovery requirements.

  • Backup ownership, schedules, retention, and locations are known.

  • Critical backups have been restored successfully.

  • Recovery procedures are documented, tested, and repeatable.

  • Actual recovery times and data-loss exposure are measured.

  • Systems without viable recovery paths are visible to leadership.

  • High-priority recovery mitigations are implemented.

  • System-owning teams understand and can maintain their recovery procedures.

  • Recovery capability does not depend entirely on one individual.

Compensation & Benefits

Base Compensation: Competitive, based on experience, portfolio strength, and geographic location.

Milestone Bonuses: 10% milestone bonus awarded to every team member assisting with the MVP build-out.

Annual Performance Bonus: Represents a significant percentage of total compensation.

Benefits for Domestic Hires: Health, dental, and vision plans.

Annual Paid Offsite: Team retreats in Caribbean, Hawaii, ski destinations, and other exciting locations.

Why Join SunCore Digital?

High-impact role in a rapidly scaling digital company.

Fully remote team with async flexibility.

Direct collaboration with executive leadership and best-in-class marketing/design partners.

Opportunity to shape Security processes and mentor junior team members.

Competitive compensation with performance upside.

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account