Cisco Logo

Cisco

Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes

Posted Yesterday
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Mid level
In-Office
London, Greater London, England, GBR
Mid level
Design, deploy, and operate highly available, multi-region SaaS infrastructure on AWS. Build Kubernetes-based platform tooling, automate deployments and operational processes, improve reliability and security, conduct scale and chaos testing, and participate in 24x7 incident response. Collaborate with software engineering teams to optimize performance, resilience, and operational excellence across a microservice platform.
The summary above was generated by AI

Please note that we operate a hybrid working model, and this role will require you to work in one of our Engineering hubs 2 days a week.

Meet the Team

Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.


ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.


Your Impact

We are seeking a skilled Site Reliability Engineer (DevSecOps) in Production Engineering with a strong background in SaaS, operations and security. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.


Responsibilities
  • Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.

Minimum Qualifications
  • Hands-on experience deploying, operating, and troubleshooting containerized workloads in production Kubernetes environments.
  • Professional experience diagnosing and administrating Linux/Unix systems, including process management, file systems, and networking protocols (TCP/IP, DNS, HTTP).
  • Professional experience developing automation tooling, operational scripts, or backend services using Python or Go.
  • Practical experience building hardened container images and integrating automated security scanning tools (SAST, DAST, or container vulnerability scanners) into CI/CD pipelines.

Preferred Qualifications
  • Familiarity with best practices for operating a large-scale, highly available enterprise platform.
  • 3+ years of experience in a related role.
  • Excellent communication and documentation skills.
  • Strong sense of ownership, drive, and attention to detail.

 

Why Cisco? 

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. 

We are Cisco, and our power starts with you. 

Similar Jobs

20 Hours Ago
Remote or Hybrid
London, Greater London, England, GBR
Mid level
Mid level
Cloud • Software
Designs, deploys, and operates highly available, secure, multi-region cloud platforms. Responsibilities include managing AWS and Kubernetes services, automating production operations, improving reliability and scalability, participating in 24x7 incident response, and developing tooling for deployment, testing, failure recovery, and platform security. The role collaborates with application engineering teams and requires Linux administration, Python or Go development, container security, and CI/CD automation experience.
Top Skills: ArgocdAWSCi/CdCncfDastDnsGoHTTPKubernetesLinuxOpentelemetryPrometheusPythonSastService MeshTcp/IpUnix
2 Months Ago
In-Office
Senior level
Senior level
Aerospace
The Senior Site Reliability Engineer oversees application tasks, designs automation for scalability, develops monitoring tools, and supports DevOps teams while collaborating with technical staff on solutions.
Top Skills: APIsAWSBashKubernetesPowershellPythonRubyTerraform
9 Hours Ago
Remote or Hybrid
Senior level
Senior level
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Business Analyst responsible for Market Risk data quality, end-to-end data lineage, process re-engineering, and AI/GenAI automation. The role analyzes data flows, change impacts, controls, and documentation while supporting FRTB and BCBS239 governance. Responsibilities include requirements gathering, process mapping, user stories, stakeholder management, and UAT coordination across Risk, Operations, Finance, and Technology teams.
Top Skills: Ai/MlAzure DevopsCopilotGenaiLlm ApisMdxPythonSQL

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account