Trayport Logo

Trayport

Principal Platform Engineer

Posted 28 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Expert/Leader
In-Office
London, Greater London, England, GBR
Expert/Leader
Designs, builds, and operates highly available multi-cloud infrastructure across AWS, Azure, Kubernetes, and on-premises environments. Owns networking, infrastructure-as-code, CI/CD, observability, disaster recovery, and SRE practices for a globally distributed trading platform. Develops AI-assisted operational workflows for incident response, automation, and toil reduction. Mentors senior engineers, serves as a cross-team design authority, and contributes to the platform roadmap in a regulated, availability-critical environment.
The summary above was generated by AI
We're hiring a Principal Platform Engineer to be a senior technical anchor in our Platform/Operations function. This is a 70% hands-on engineering role: you'll design, build, and operate the infrastructure that keeps a global trading platform running, while acting as a technical mentor and design authority across the wider team.
A core part of this role is leading our adoption of AI within operations — using LLM-based tooling, agentic workflows, and automation to reduce toil, accelerate incident response, and raise the operational leverage of the whole department. We're not looking for someone to write a strategy deck about AI; we're looking for someone who will build with it.What you'll do

Platform & cloud engineering (the 70%)

  • Design, build, and operate infrastructure across AWS (VPC, networking, EKS), Azure (AKS, AKV, networking), with an understand of on-prem environments

  • Own core networking design and implementation across within Cloud environments — connectivity, routing, DNS, load balancing, firewalls, and private links between cloud and on-prem

  • Engineer for high availability: multi-region and multi-AZ architectures, failover design, capacity planning, and disaster recovery for a platform where downtime has direct market impact

  • Build and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes platforms as products consumed by engineering teams

SRE & reliability

  • Define and drive SLOs, error budgets, and observability standards (metrics, logging, tracing) across the platform

  • Take part in post-incident reviews and help with the resulting reliability work to completion

  • Continuously reduce toil through automation — if we've done it manually twice, you're already scripting it

AI-driven operations

  • Identify, prototype, and productionise AI-assisted workflows across the department: incident triage and summarisation, runbook automation, log/alert analysis, change-risk assessment, internal knowledge tooling

  • Use AI-assisted engineering tools (Copilot, or equivalents) as a first-class part of your own workflow, and coach the team to do the same safely and effectively

  • Establish sensible guardrails for AI use in a regulated, availability-critical environment, knowing when automation should act and when it should recommend

Technical leadership (the 30%)

  • Act as a mentor to engineers across the Platform and Operations teams, raising the bar on design, code, and operational practice

  • Be a design authority on cross-team projects: review architectures, challenge assumptions, and ensure new services are built to be operable, observable, and resilient from day one

  • Contribute to the technical roadmap for the platform function, balancing reliability investment against delivery

What we're looking for

Must have

  • A background in Operations or SRE running highly available, redundant production platforms — you understand failure domains, graceful degradation, and what "five nines" costs

  • Deep hands-on experience with AWS (VPC design, networking, EKS) and Azure (AKS, Key Vault, networking) — genuinely multi-cloud, not one cloud plus a certification

  • Strong Networking fundamentals: TCP/IP, routing concepts, firewalls, load balancing, hybrid connectivity (Direct Connect / ExpressRoute, VPNs)

  • Production Kubernetes experience at scale, including day-2 operations (upgrades, capacity, security, multi-cluster)

  • Infrastructure-as-code and automation as a default working style (Terraform, Ansible, or similar; strong scripting in Python, Go, or Bash)

  • Demonstrable, practical use of AI tooling to improve engineering or operational workflows — you can show us something you've automated, accelerated, or de-toiled with it

  • The credibility and communication skills to mentor senior engineers and influence design decisions without formal authority

  • Understanding of different database technologies 

Nice to have

  • Experience in trading, exchanges, market data, fintech, or another latency- and availability-sensitive domain

Trayport is committed to creating and sustaining a collegial work environment in which all individuals are treated with dignity and respect and one which reflects the diversity of the community in which we operate. We provide accommodations for applicants and employees who require it.

HQ

Trayport London, England Office

7th Floor, 9 Appold Street, London, United Kingdom, EC2A 2AP

Similar Jobs

4 Hours Ago
Hybrid
Staines-upon-Thames, Middlesex, England, GBR
Expert/Leader
Expert/Leader
Information Technology • Software
Build and operate shared authentication and authorization platform capabilities across cloud, legacy, and lifecycle environments. Responsibilities include developing platform services, APIs, infrastructure automation, self-service workflows, and Go-based tooling; operating Kubernetes and containerized services; implementing OAuth 2.0, OpenID Connect, token management, permission models, tenant isolation, and enterprise federation; and ensuring reliability through observability, testing, safe rollouts, rollback, backup, and recovery. The role also provides technical leadership through architecture, code reviews, standards, and production ownership.
Top Skills: Ci/CdCloud-Native TechnologiesContainersCurityGitopsGoInfrastructure As CodeKafkaKeycloakKubernetesLlmsOauth 2.0Openid ConnectPostgresRedpandaSAMLSpicedb
21 Days Ago
In-Office
London, Greater London, England, GBR
Expert/Leader
Expert/Leader
Fintech
Leads the architecture, operation, and evolution of global multi-cloud platforms across GCP and AWS. Builds AWS landing zones, self-service developer platforms, CI/CD tooling, observability, automation, governance, and security capabilities. Supports enterprise AI transformation through platform services, guardrails, and operational patterns for AI applications and agentic workflows. Provides technical leadership during incidents, evaluates emerging technologies, mentors engineers, and drives reliability, cost efficiency, developer experience, and engineering standards across distributed teams.
Top Skills: AirflowApigeeAPIsAWSCi/CdDatadogElkGCPGithub ActionsGrafanaInfrastructure As CodeKubernetesPrometheusPythonRetrieval-Augmented Generation (Rag)Serverless TechnologiesSplunkTerraformVector Databases
24 Days Ago
In-Office
Entry level
Entry level
Fintech
Builds and operates secure, scalable infrastructure platforms for enterprise agentic AI adoption. Responsibilities include developing Terraform modules, deployment automation, CI/CD pipelines, observability, identity and security-as-code capabilities, documentation, and self-service deployment experiences. The role translates architectures into tested platform solutions, supports shared services, troubleshoots consuming-team issues, establishes engineering standards, drives roadmap adoption, and leads complex infrastructure and AI enablement decisions across architecture, security, operations, and application teams.
Top Skills: Amazon BedrockAzure Ai FoundryCi/CdCloud Identity And Access ManagementCloud NetworkingEvaluation FrameworksInfrastructure As CodeModel GatewaysObservabilitySecurity As CodeTerraformVector Databases

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account