Carbon3.ai Logo

Carbon3.ai

Site Reliability Engineer

Reposted 9 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Build AI-driven SRE tooling and agentic automation to triage, diagnose, and remediate infrastructure issues. Integrate LLM-powered agents with observability, ITSM, and infrastructure APIs, develop self-service tooling and ChatOps, tune event/alert intelligence, and convert runbooks into auditable executable automations while contributing to operational standards and incident learning.
The summary above was generated by AI

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations



Role Summary:

We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

 

This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response.

 

Key Responsibilities:

  • Build Python-based automation for incident triage, runbook execution, and routine operational tasks.
  • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
  • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
  • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
  • Maintain version-controlled runbook-as-code and automation libraries.
  • Translate post-incident learnings into better tooling, automation and operational standards.
  • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

 

Essential Experience:

  • Experience in SRE, Platform Engineering, or production infrastructure operations.
  • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
  • Exposure to incident management / on-call and converting manual runbooks into automation.
  • Experience with Python for automation, APIs and integrations.

 

Nice To Have:

  • GPU, datacentre or colocation infrastructure experience.
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
  • ChatOps tooling (Slack or Microsoft Teams bots).
  • OpenTelemetry, logging or distributed tracing experience.
  • DCIM, IPAM or hypervisor-control-plane integrations.
  • Experience with LLM-assisted or agent-based operational automation.

 

Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

 

Diversity & Inclusion:

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

 

HQ

Carbon3.ai Tonbridge and Malling, England Office

Tonbridge and Malling, United Kingdom

Similar Jobs

8 Days Ago
Easy Apply
Remote
United Kingdom
Easy Apply
Expert/Leader
Expert/Leader
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills: Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
3 Hours Ago
Remote or Hybrid
2 Locations
Senior level
Senior level
Retail • Software
Leads the team building the SRE function and owns its technical vision, platform capabilities, reusable tooling, pipelines, frameworks, and developer experience. Manages senior engineering staff, drives recruitment and retention, aligns technical strategy, oversees reliability and operational excellence, partners with vendors, and contributes hands-on code. The role requires strong SRE, cloud, software architecture, DevOps, testing, reliability engineering, and people leadership experience.
Top Skills: AutomationCapacity PlanningCi/CdCloud InfrastructureContainersDistributed SystemsError BudgetsIncident ResponseKubernetesObservabilityPerformance EngineeringProgressive DeliverySelf-Healing SystemsSlisSlos
3 Days Ago
Remote or Hybrid
United Kingdom
Expert/Leader
Expert/Leader
Financial Services
Leads the design and operation of highly available, scalable, and observable production infrastructure. Defines SLOs, error budgets, and reliability targets; drives incident response, root-cause analysis, and postmortems; develops automation and deployment tooling; champions monitoring and observability; evaluates platform technologies; and mentors engineers. The role also applies secure AI-assisted engineering practices to improve incident triage, testing, and delivery workflows.
Top Skills: AWSBashGoJavaKubernetesPythonTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account