Autodesk Logo

Autodesk

Senior Site Reliability Engineer

Posted 3 Days Ago
In-Office or Remote
Hiring Remotely in Greece
Senior level
In-Office or Remote
Hiring Remotely in Greece
Senior level
Lead reliability automation projects and build shared tooling, platform services, observability, self-healing, incident response, resilience testing, and disaster recovery capabilities. Partner across engineering, infrastructure, security, and operations teams to improve service reliability, scalability, deployment safety, and operational efficiency. Own production services and participate in a 24x7 on-call rotation while eliminating operational toil through software engineering.
The summary above was generated by AI

Job Requisition ID #


26WD101322

Senior Site Reliability Engineer, Reliability Automation


Position Overview


Autodesk is the global leader in design and make technology, including industry-leading 3D design, engineering, and entertainment software and services, that offer customers better outcomes through automation and insights for their design and make processes. If you’ve ever driven a high-performance car, admired a towering skyscraper, used a smartphone, or watched a great film, chances are you’ve experienced what millions of Autodesk customers are doing with our software.


Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk products and customers.


As part of an SRE team focused on Reliability Automation, you will have a unique opportunity to help shape how Autodesk detects, responds to, and engineers away production issues at scale. This is a project-oriented, engineering-heavy role where you will build the automation, tooling, and platform capabilities that other engineering teams rely on to run their services reliably.


You will combine software engineering and production operations to automate how Autodesk services are operated, monitored, recovered, and validated. You will partner closely with product engineering, platform, infrastructure, and operations teams to turn manual operational work into durable, reusable automation.


The ideal candidate is a strong software engineer who is drawn to operational problems. You should be comfortable designing and shipping production-grade software, and equally comfortable reasoning about how distributed systems fail. Success in this role requires strong technical judgment, the ability to drive projects end to end, and a passion for replacing manual operational work with durable engineering.


Responsibilities


  • Lead reliability automation projects end to end, from problem definition and design through delivery, adoption, and measurable impact on availability, performance, and operational efficiency.
  • Design, develop, and maintain shared automation, tooling, and platform services that improve the reliability, scalability, and operability of production systems across Autodesk.
  • Partner with engineering teams to understand their operational pain, and drive adoption of the automation and reliability capabilities your team builds.
  • Instrument and automate reliability signals such as SLOs/SLIs, error budgets, and alerting-as-code, so teams get reliability measurement by default rather than by manual effort.
  • Build automation that improves deployment safety, operational efficiency, and service recovery, including self-healing and auto-remediation capabilities.
  • Automate incident response workflows, diagnostics, and runbook execution to reduce time to detect, time to mitigate, and manual intervention.
  • Extend monitoring, alerting, logging, and tracing capabilities, and automate their coverage and consistency across services.
  • Own the reliability, availability, and operability of the automation and platform services your team runs, including on-call response when they fail.
  • Turn incident and post-incident findings from across the organization into automation, tooling, and durable engineering fixes.
  • Design and scale automated resilience testing, chaos engineering, Gameday, and disaster recovery tooling to validate system behavior and recovery capabilities.
  • Continuously identify and eliminate operational toil through software engineering, automation, and process improvement.
  • Ensure the automation and services your team builds meet Autodesk security, privacy, and operational risk requirements.
  • Participate in a 24x7 on-call rotation for the automation, tooling, and platform services owned by the team.
  • Function effectively in a fast-paced environment while helping mature reliability automation practices across engineering.

Basic Qualifications

  • B.S. or higher in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
  • 7+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, or a closely related discipline.
  • Strong software engineering fundamentals, including API and service design, testing, code review, and shipping production-grade software.
  • Working knowledge of reliability engineering concepts such as SLOs/SLIs, observability, incident response, and toil reduction, and experience applying them in practice.
  • Experience building on AWS, Azure, or another public cloud platform.
  • Strong programming skills in languages such as Python, Go, Java, or similar.
  • Experience with Infrastructure as Code, CI/CD pipelines, and deployment automation.
  • Ability to work independently, scope ambiguous problems, and drive projects to completion across team boundaries.
  • Strong written and verbal communication skills.

Preferred Qualifications


  • 7+ years of experience building software and automation for production systems.
  • Experience building self-healing, auto-remediation, or event-driven operational automation.
  • Experience building internal platforms, developer tooling, or services adopted by other engineering teams.
  • Experience building or adopting AI-assisted operational tooling, such as LLM-based incident triage, diagnostics, or agentic remediation workflows.
  • Experience with containers, Kubernetes, cloud-native architectures, APIs, and distributed systems.
  • Experience integrating with observability platforms such as Splunk, Dynatrace, Datadog, CloudWatch, or similar.
  • Experience with event-driven architectures, workflow orchestration, or job scheduling systems.
  • Experience with incident management platforms such as incident.io, FireHydrant, or PagerDuty, and automating workflows on top of them.
  • Experience designing and implementing operational automation at scale.
  • Experience building or automating Gamedays, chaos experiments, disaster recovery exercises, or resilience testing.
  • Experience using incident and operational data to identify, prioritize, and measure automation opportunities.
  • Strong collaboration skills and ability to work effectively across engineering, security, and operations teams.
  • Passion for building reliable, secure, and scalable systems that customers can trust.

#LI-MM1

Learn More


About Autodesk

Welcome to Autodesk! Amazing things are created every day with our software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies. We help innovators turn their ideas into reality, transforming not only how things are made, but what can be made.


We take great pride in our culture here at Autodesk – it’s at the core of everything we do. Our culture guides the way we work and treat each other, informs how we connect with customers and partners, and defines how we show up in the world.


When you’re an Autodesker, you can do meaningful work that helps build a better world designed and made for all. Ready to shape the world and your future? Join us!


Salary transparency

Salary is one part of Autodesk’s competitive compensation package. For Croatia based roles, we expect a starting base salary between 38,000 EUR and 55,000 EUR. Offers are based on the candidate’s experience and geographic location, and may exceed this range. In addition to base salaries, our compensation package may include annual cash bonuses, commissions for sales roles, stock grants, and a comprehensive benefits package.

Belonging
We take pride in cultivating a culture of belonging where everyone can thrive. Learn more here: https://www.autodesk.com/company/global-belonging


In-Person Onboarding and Identity Verification

This role may require in-person onboarding and/or in-person ID verification.

Autodesk London, England Office

London, United Kingdom

Similar Jobs

10 Days Ago
Remote
Senior level
Senior level
Information Technology • Software
Operate and scale distributed bare-metal GPU and virtualized infrastructure across Linux, MAAS, Kubernetes, networking, observability, and multiple sites. Automate provisioning and operations, manage virtualization and hardware, define reliability practices, lead incident response, maintain on-call coverage, and build internal infrastructure tooling. The role also owns lifecycle management, security hardening, runbooks, and cross-functional improvements to infrastructure reliability and efficiency.
Top Skills: AlertmanagerAnsibleAtlantisBashCephCloud-InitCloudflare ApisCniDebianDnsFirewallsGitGoGrafanaIpmiKubernetesKvm/LibvirtL2/L3 RoutingLinuxMaasNetboxOpenstackOpentofuPrometheusProxmoxPxePythonRaidRbacRedfishService MeshSopsTerraformUbuntuUnifiVaultVictorialogsVictoriametricsVlanVMwareVpnZfs
11 Days Ago
In-Office or Remote
Senior level
Senior level
Fitness • Healthtech
Operate and improve AWS infrastructure, Kubernetes platforms, infrastructure-as-code, CI/CD, observability, and internal tooling. Participate in on-call rotations, investigate and resolve complex production incidents, automate recurring operational work using Go, and collaborate across engineering, product, and security teams. The role requires hands-on ownership of production systems, structured debugging, rapid learning, and strong technical communication.
Top Skills: Amazon EksArgocdAtlantisAWSCi/CdCiliumGitopsGoGrafanaInfrastructure As CodeIstioKarpenterKubernetesLokiPrometheusTempoTerraformTerramate
15 Days Ago
Remote
Senior level
Senior level
Cloud • Security • Software • Generative AI
Design, build, scale, and mature Elastic’s multi-cloud platform infrastructure. Develop software, tooling, and automation to improve reliability, scalability, and operational excellence. Lead initiatives, prevent recurring customer impact, support incident and problem management, and participate in a follow-the-sun on-call rotation. Collaborate across distributed teams while mentoring and uplifting engineers.
Top Skills: CrossplaneDockerElastic StackGoGraphiteInfluxKubernetesLinuxPrometheusPublic CloudTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account