JPMorganChase Logo

JPMorganChase

Lead Site Reliability Engineer

Reposted One Month Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Senior level
Hybrid
London, Greater London, England, GBR
Senior level
Lead SRE embedded with front-office trading teams to improve reliability, observability, and performance across low-latency, global trading platforms. Responsibilities include incident response, root cause analysis, coding reliability improvements (Java/Kotlin/Python), designing SRE patterns (automation, self-healing), observability, and partnering with infrastructure, cloud, and cybersecurity teams.
The summary above was generated by AI

Our trading technology stack is undergoing a multi‑year convergence and modernization journey. You will play a pivotal role in shaping our next‑generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‑paced front‑office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization.

As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management’s Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front‑office trading platforms.


Job Responsibilities:

  • Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions.
  • Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. 
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability.

Required Qualifications, Capabilities, and Skills:

  • Strong hands on experience in front office trading environments or similarly high pressure, low latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission critical systems.
  • Proven ability to lead incident response and drive long term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures.
  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations.
  • Strong leadership presence with a collaborative mindset.
About UsJ.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.
  
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
About the TeamJ.P. Morgan Asset & Wealth Management delivers industry-leading investment management and private banking solutions. Asset Management provides individuals, advisors and institutions with strategies and expertise that span the full spectrum of asset classes through our global network of investment professionals. Wealth Management helps individuals, families and foundations take a more intentional approach to their wealth or finances to better define, focus and realize their goals.​

JPMorganChase London, England Office

25 Bank Street, Canary Wharf, London, United Kingdom, E14 5JP

Similar Jobs

21 Days Ago
Remote or Hybrid
United Kingdom
Expert/Leader
Expert/Leader
Financial Services
Leads the design and operation of highly available, scalable, and observable production infrastructure. Defines SLOs, error budgets, and reliability targets; drives incident response, root-cause analysis, and postmortems; develops automation and deployment tooling; champions monitoring and observability; evaluates platform technologies; and mentors engineers. The role also applies secure AI-assisted engineering practices to improve incident triage, testing, and delivery workflows.
Top Skills: AWSBashGoJavaKubernetesPythonTerraform
12 Days Ago
In-Office
London, Greater London, England, GBR
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Lead the architecture and development of an AI-powered Production Intelligence platform combining SRE, observability, agentic AI, and automated remediation. Build AI agents, MCP integrations, telemetry pipelines, anomaly detection, evaluation frameworks, and safe human-in-the-loop workflows. Define reliability practices including SLIs, SLOs, error budgets, incident response, and post-incident reviews. Mentor engineers, establish technical standards, and coordinate delivery across global application and infrastructure teams.
Top Skills: APIsAWSAzureC++DockerElasticsearchGCPGithub ActionsGitopsGoGrafanaJavaJenkinsKafkaKubernetesKubernetesLlmsMicroservicesModel Context ProtocolOpentelemetryPostgresPrometheusPythonRagRedisSplunkSreTerraformThousandeyesVector Databases
22 Days Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Financial Services
Lead SRE responsible for driving reliability practices, designing resilient systems, conducting resiliency reviews, leading incident response, mentoring engineers, improving observability and CI/CD, and adopting validated AI-assisted reliability workflows across the SDLC.
Top Skills: AWSDatadogDockerDynatraceEcsGitlabGrafanaJenkinsKubernetesPrometheusPythonSplunkTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account