The Depository Trust & Clearing Corporation (DTCC) Logo

The Depository Trust & Clearing Corporation (DTCC)

IAM Engineering Director

Posted 6 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Expert/Leader
In-Office
London, Greater London, England, GBR
Expert/Leader
Lead IAM Site Reliability Engineering to ensure availability, scalability, observability, and resiliency of IAM platforms (PAM, AD, PKI, secrets, auth, cloud identity). Define observability standards, monitoring architecture, SLIs/SLOs, incident management, resilience testing, automation, and partner with engineering and infrastructure teams to embed reliability-by-design.
The summary above was generated by AI

Are you ready to make an impact at DTCC?  

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets.  We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact.  We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

Pay and Benefits:

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits
  • Pension
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (Onsite Tuesdays, Wednesdays and a third day of your choosing)

The impact you will have in this role:

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.  
The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Role Summary:

We are seeking an experienced Director of IAM Site Reliability Engineering (SRE) & Observability to lead the reliability, availability, operational architecture, monitoring strategy, and resiliency of the enterprise Identity and Access Management (IAM) ecosystem.

This leader will be responsible for ensuring the health, performance, scalability, recoverability, and operational readiness of critical IAM platforms including Privileged Access Management (PAM), Active Directory, Certificate Management Infrastructure (PKI), Secrets Management, Authentication Services, Cloud Identity Platforms, and related Identity Security services.

The role will establish enterprise-wide observability standards, develop service reliability architectures, define monitoring frameworks, and partner closely with Engineering teams to ensure IAM platforms are designed, instrumented, and operated for maximum availability and resilience.

The successful candidate will drive a proactive reliability culture focused on monitoring, telemetry, automation, service health, incident prevention, operational architecture, and continuous improvement.

Your Primary Responsibilities:

IAM Site Reliability Engineering Leadership

  • Lead the IAM Site Reliability Engineering (SRE) function across all IAM platforms and services.
  • Own platform availability, service health, resiliency, observability, and operational readiness objectives.
  • Establish reliability engineering practices and operational excellence standards across the IAM ecosystem.
  • Partner with engineering teams throughout the software and platform lifecycle to embed reliability-by-design principles.

Observability & Monitoring Strategy

  • Define and implement enterprise observability standards across all IAM platforms.
  • Develop monitoring architecture standards covering: 
    • Infrastructure Monitoring
    • Application Monitoring
    • User Experience Monitoring
    • Transaction Monitoring
    • Dependency Monitoring
    • Security Event Monitoring
    • Cloud Service Monitoring
  • Establish standards for logging, metrics, tracing, dashboards, alerting, correlation, and telemetry collection.
  • Drive implementation of centralized observability platforms leveraging tools such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, App Insights, or equivalent solutions.
  • Define platform-specific monitoring requirements and operational instrumentation standards for all IAM services.
  • Ensure all IAM applications meet monitoring, alerting, and observability requirements prior to production deployment.

Reliability Architecture & Engineering

  • Create operational architecture diagrams, service dependency maps, data flow diagrams, and platform resiliency models for IAM services.
  • Work closely with IAM Engineering teams to review solution designs and identify availability, scalability, and resiliency risks.
  • Establish architecture standards for: 
    • High Availability (HA)
    • Disaster Recovery (DR)
    • Multi-site Resiliency
    • Failover Design
    • Capacity Planning
    • Fault Tolerance
    • Service Recovery
  • Conduct reliability design reviews and production readiness assessments for IAM platforms and applications.

Availability & Resilience Management

  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and Availability Targets.
  • Drive initiatives to improve platform uptime, stability, recoverability, and performance.
  • Identify and eliminate single points of failure and operational bottlenecks.
  • Lead resiliency testing exercises, failover validation, disaster recovery testing, and business continuity preparedness.

Incident & Service Health Management

  • Lead major incident management, executive communications, and post-incident reviews.
  • Analyze incident trends, recurring failures, and reliability gaps to drive platform improvements.
  • Establish proactive service health review processes and operational risk assessments.
  • Ensure corrective actions are tracked and implemented to prevent repeat incidents.

Automation & Operational Excellence

  • Drive automation of monitoring, alert triage, remediation, service validation, and operational workflows.
  • Promote Infrastructure as Code (IaC), automated health checks, self-healing capabilities, and operational engineering practices.
  • Improve alert quality by reducing noise and focusing on actionable service indicators.
  • Establish reliability scorecards and operational maturity frameworks across IAM platforms.

Cross-Functional Leadership

  • Partner with IAM Engineering, Security Engineering, Enterprise Architecture, Cloud, Infrastructure, Network, and Vendor teams.
  • Serve as the primary authority for IAM platform observability, availability, and operational architecture standards.
  • Influence engineering roadmaps by incorporating reliability, monitoring, and resiliency requirements into platform design and development processes.

**NOTE:  The Primary Responsibilities of this role are not limited to the details above. **

Qualifications:

  • Bachelor's degree preferred or equivalent experience

Talents Needed For Success:

  • Minimum of 10 years related experience
  • 12+ years of experience in Site Reliability Engineering, Platform Engineering, IAM, Infrastructure Engineering, or Cybersecurity.
  • Experience supporting complex IAM environments including PAM, Active Directory, PKI, Authentication Services, Secrets Management, and Cloud Identity platforms.
  • Strong expertise in observability, monitoring architecture, telemetry design, and platform instrumentation.
  • Experience creating architecture diagrams, dependency maps, operational blueprints, and resiliency designs.
  • Hands-on experience with enterprise monitoring and observability platforms.
  • Strong knowledge of High Availability, Disaster Recovery, Business Continuity, and resilience engineering.
  • Experience leading major incident management and operational transformation programs.
  • Proven ability to influence architecture and engineering teams without direct ownership.
  • Strong executive communication and stakeholder management skills.

We offer top class training and development for you to be an asset in our organization!

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.


The Depository Trust & Clearing Corporation (DTCC) London, England Office

London, United Kingdom

Similar Jobs

30 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Own Perk’s US and UK PR strategy, building Tier 1 relationships with business and financial media. Lead proactive pitching, crisis communications, corporate narrative development, executive profiling, media training, and translation of AI and technology developments into compelling stories. Manage PR agencies, measure coverage and ROI, and integrate AI tools into research, drafting, monitoring, and reporting. Partner closely with senior leadership, product, engineering, and global communications teams in a fast-paced, hybrid environment.
Top Skills: Artificial Intelligence (Ai)SaaS
38 Minutes Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain resilient systems, automation, operational APIs, observability, and infrastructure-as-code tooling. Diagnose distributed-system incidents from edge to origin, manage monitoring and alerting platforms, configure Cloudflare services, participate in incident response and post-mortems, and improve reliability, performance, and operational consistency. Collaborate across SRE, development, and IT Operations while mentoring colleagues and applying AI tools to increase productivity and root-cause analysis.
Top Skills: AnsibleCdnCloudflareCoding AssistantsDdos ProtectionDnsGoGrafanaInfrastructure As CodeJavaScriptLlm PlatformsNew RelicOpentelemetryPagerdutyPythonShell ScriptingSplunkTerraformWaf
44 Minutes Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Leads the re-architecture and delivery of business-critical risk and regulatory platforms using Golang, React, cloud technologies, Kafka, SQL, .NET, and TypeScript. Provides technical leadership, governance, solution design, code-quality oversight, estimation, documentation, escalation support, and mentoring. Ensures highly available, scalable, maintainable, performant, and cross-browser-compatible systems while collaborating with stakeholders and architects from development through production.
Top Skills: .NetCloud PlatformsGoKafkaReactSQLTypescript

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account