Carbon3.ai Logo

Carbon3.ai

Site Reliability Engineer

Reposted 22 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Senior level
Remote
Hiring Remotely in United Kingdom
Senior level
Build AI-driven SRE tooling and agentic automation to triage, diagnose, and remediate infrastructure issues. Integrate LLM-powered agents with observability, ITSM, and infrastructure APIs, develop self-service tooling and ChatOps, tune event/alert intelligence, and convert runbooks into auditable executable automations while contributing to operational standards and incident learning.
The summary above was generated by AI

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations



Role Summary:

We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

 

This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response.

 

Key Responsibilities:

  • Build Python-based automation for incident triage, runbook execution, and routine operational tasks.
  • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
  • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
  • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
  • Maintain version-controlled runbook-as-code and automation libraries.
  • Translate post-incident learnings into better tooling, automation and operational standards.
  • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

 

Essential Experience:

  • Experience in SRE, Platform Engineering, or production infrastructure operations.
  • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
  • Exposure to incident management / on-call and converting manual runbooks into automation.
  • Experience with Python for automation, APIs and integrations.

 

Nice To Have:

  • GPU, datacentre or colocation infrastructure experience.
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
  • ChatOps tooling (Slack or Microsoft Teams bots).
  • OpenTelemetry, logging or distributed tracing experience.
  • DCIM, IPAM or hypervisor-control-plane integrations.
  • Experience with LLM-assisted or agent-based operational automation.

 

Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

 

Diversity & Inclusion:

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

 

HQ

Carbon3.ai Tonbridge and Malling, England Office

Tonbridge and Malling, United Kingdom

Similar Jobs

9 Days Ago
Remote or Hybrid
London, Greater London, England, GBR
Mid level
Mid level
Cloud • Software
Design, operate, and scale large distributed systems for telemetry processing. Build automation, use AI tooling to reduce toil, ensure availability and disaster recovery, participate in on-call incident response, troubleshoot production AWS/Kubernetes environments, and collaborate with application teams to meet SLOs/SLAs.
Top Skills: Ai ToolingAWSGnu/LinuxGoKubernetesPythonTerraform
2 Days Ago
Remote
United Kingdom
Entry level
Entry level
Information Technology
Own reliability and observability for critical product journeys in a distributed consumer mobile product. Build metrics, dashboards, alerts, SLIs, and SLOs; improve monitoring, logging, tracing, and incident response; act as a first responder for P0/P1 incidents; investigate and mitigate production issues; coordinate escalations; lead postmortems; and drive infrastructure and tooling improvements using Node.js, TypeScript, and AWS.
Top Skills: AWSCloudflareCloudwatchNode.jsReact NativeSentryTypescript
3 Days Ago
Remote or Hybrid
United Kingdom
Entry level
Entry level
Utilities
Leads the design, development, and operation of reliable production systems. Drives improvements in observability, alerting, incident management, operational readiness, scalability, security, and maintainability. Supports teams during incidents, guides testing and architecture decisions, promotes learning from incidents, and mentors engineers across the organization. Partners with product teams to make customer-focused, data-informed, and cost-efficient technical decisions while advancing agile delivery practices.
Top Skills: AgileAIAlertingCloud-Native ArchitectureIncident ManagementObservabilitySaaSSecure Coding

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account