WebMD Logo

WebMD

Site Reliability Engineer, K8s (Remote International)

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United Kingdom
Mid level
Remote
Hiring Remotely in United Kingdom
Mid level
Design, build, and operate a large-scale Kubernetes platform supporting distributed systems, data infrastructure, and business-critical services. Responsibilities include Kubernetes lifecycle management, reliability and observability, incident response, infrastructure automation, GitOps workflows, networking, platform security, developer self-service, and bare-metal operations. The role requires hands-on architecture, operational excellence, and technical leadership across infrastructure platforms.
The summary above was generated by AI
WebMD and its affiliates is an Equal Opportunity/Affirmative Action employer and does not discriminate on the basis of race, ancestry, color, religion, sex, gender, age, marital status, sexual orientation, gender identity, national origin, medical condition, disability, veterans status, or any other basis protected by law. 
Build the platform behind PulsePoint

PulsePoint operates large-scale data and advertising platforms that power business-critical services used every day across the company.

Our Platform Engineering team owns the foundation that enables engineering teams to move quickly and safely. We build and operate the infrastructure that supports Kubernetes workloads, data platforms, developer tooling, observability and production operations at scale.

Unlike many cloud-only environments, we own the full lifecycle of the infrastructure - from bare-metal hardware and networking to Kubernetes, observability and developer experience.

The environment supports large-scale Kubernetes workloads, multi-petabyte data systems and business-critical services used across multiple engineering organizations.

We're looking for an experienced engineer to help shape its next stage of growth.

This is a hands-on role focused on architecture, reliability, automation and operational excellence. You will work on complex distributed systems, drive platform improvements and help shape the technical direction of infrastructure across the company.

What you'll work on

You'll help design, build and operate the Kubernetes platform used across PulsePoint.

Examples of challenges you may work on include:

  • Platform architecture and Kubernetes lifecycle management

  • Reliability, observability and incident response

  • Infrastructure automation and GitOps workflows

  • Networking, service connectivity and platform security

  • Developer experience and self-service platform capabilities

  • Large-scale distributed systems running on bare-metal infrastructure

Technology

You'll work in an environment that includes:

  • Kubernetes and platform services

  • Multi-petabyte data infrastructure

  • Bare-metal and cloud environments

  • GitOps and infrastructure automation

  • Modern observability and reliability engineering practices

Technologies commonly used across the environment include Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis and Ceph.

Experience with every technology is not required.

Who we’re looking for

Success in this role is not measured by the number of tickets closed or clusters operated. Success means building platform capabilities that make engineering teams more reliable, productive and autonomous.

We're especially interested in engineers who:

  • Have operated production infrastructure at meaningful scale

  • Understand how distributed systems fail and recover

  • Prefer automation over repetitive operational work

  • Enjoy simplifying systems rather than adding complexity

  • Take ownership beyond the boundaries of a single component

  • Willing and able to work 9am-6pm ET U.S. hours. You can work fully remotely
Why this role

This is not a ticket-driven operational role.

You'll help define platform architecture, influence engineering standards and work on infrastructure that supports multiple engineering organizations.

The engineer joining this role is expected to become a key technical contributor shaping the future of the platform.

We try to keep the process focused and practical.

  1. Introductory conversation (~60m)
    Learn about your background and discuss the role.

  2. Technical discussion (~60m)
    Deep dive into systems engineering, Kubernetes and operational experience.

  3. Architecture discussion (~60m)
    Explore platform design, distributed systems and technical decision making.

  4. Leadership conversation (~30m)
    Meet engineering leadership and discuss team, strategy and long-term direction.

Similar Jobs

23 Days Ago
Remote
United Kingdom
Mid level
Mid level
Social Media
The Site Reliability Engineer will design, build, and maintain AWS cloud infrastructure, ensure performance and reliability, automate tasks, and participate in incident management.
Top Skills: AWSBashPythonTerraform
21 Minutes Ago
Remote or Hybrid
2 Locations
Senior level
Senior level
Financial Services
Supports strategy, governance, roadmap development, commercialization, and operational execution for J.P. Morgan’s EMEA Account Solutions portfolio. Partners with product, finance, treasury, operations, technology, legal, compliance, risk, sales, and implementation teams. Defines requirements and customer journeys, coordinates strategic initiatives, analyzes product and market data, prepares executive communications, maintains risk controls, and supports regulatory changes and client-focused product enhancements.
Top Skills: ExcelMicrosoft Powerpoint
2 Hours Ago
Remote or Hybrid
2 Locations
Entry level
Entry level
Financial Services
Design and deliver enterprise-grade machine learning systems and AI-powered applications for banking. Build scalable, secure, distributed applications and automated cloud, desktop, and ML pipelines. Apply MLOps practices for versioning, reproducibility, and observability while collaborating with cloud and site reliability engineering teams. Develop reusable libraries and business-critical, data-intensive solutions, aligning machine learning problems with business objectives. The role may include mentoring or people leadership.
Top Skills: AWSCloud ComputingDatabasesDistributed SystemsKubernetesLanggraphMachine LearningMessaging And Queue SystemsMlopsMulti-ThreadingPydantic AiPythonTypescript

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account