Lightning AI

London
50 Total Employees
Year Founded: 2019

Jobs at Lightning AI

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

2 Days AgoSaved
Hybrid
3 Locations
Artificial Intelligence • Machine Learning • Software
Design and build backend systems for AI agent orchestration, tool execution, workflow management, memory, and state handling. Develop scalable APIs and reliable infrastructure for distributed AI workflows, improve platform reliability and observability, and partner with product, research, and engineering teams to deliver agent capabilities. Own technical initiatives from architecture through production, maintain software quality through testing and continuous delivery, and mentor engineers on backend and distributed-systems practices.
2 Days AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Develop and scale Lightning AI’s backend platform using Go, spanning APIs, infrastructure, billing, security, and integrations. Own features end to end, design reliable and scalable systems, improve architecture and performance, maintain software quality and continuous delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a fast-changing SaaS environment.
2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
Artificial Intelligence • Machine Learning • Software
Support ML engineering teams running production training and inference workloads across Kubernetes, cloud, Linux, and GPU infrastructure. Diagnose distributed PyTorch, CUDA, NCCL, networking, storage, scheduling, and performance issues; analyze observability data; advise customers during incidents; and improve reliability through tooling, automation, documentation, runbooks, and operational processes. The role is hybrid in London, requires at least two office days weekly, and supports EMEA shifts.
2 Days AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build and deploy production AI systems for customers, translating business objectives into scalable technical solutions. Responsibilities include architecture, proof-of-concepts, software development, deployment, monitoring, debugging, inference optimization, and distributed systems operation. The engineer partners with customer engineering teams, collaborates with product and engineering, improves reusable platform capabilities, and owns technical engagements from discovery through production scaling.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
This is a general talent community opportunity rather than a specific open position. Lightning AI invites candidates to submit their information for consideration as future roles become available across U.S. and London hubs, with occasional remote opportunities. The company develops tools and infrastructure for building, training, and deploying AI systems and values ownership, urgency, communication, teamwork, continuous improvement, and long-term thinking.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Operate, scale, and optimize distributed storage infrastructure supporting large-scale AI/ML and HPC workloads. Build Python automation, manage Linux bare-metal systems, troubleshoot storage, hardware, networking, and operating system issues, and improve performance, reliability, monitoring, capacity planning, and lifecycle management. Collaborate with infrastructure, networking, platform, and data center teams on storage deployments and scaling strategies.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Build, validate, and operate large-scale bare-metal GPU infrastructure for AI/ML and HPC workloads. Responsibilities include managing Linux systems, image pipelines, test clusters, provisioning, firmware and driver validation, GPU diagnostics, performance analysis with NVIDIA DCGM, automation, virtualization, and hardware management interfaces. The role requires troubleshooting across hardware and software layers while collaborating with infrastructure, hardware, data center, platform, and ML teams.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
2 Days AgoSaved
Remote or Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Develop and post-train deep learning models while building systems, tooling, and workflows for training, evaluation, debugging, and deployment. Contribute to open-source projects, distributed AI infrastructure, backend services, and developer platforms. Collaborate with research, product, infrastructure, and customer teams to solve complex technical problems, prototype ideas, and productionize successful experiments.
Artificial Intelligence • Machine Learning • Software
Design and scale backend systems for AI agent orchestration, distributed workflows, experiment management, platform APIs, and developer tooling. Build reliable cloud-native infrastructure, improve observability and performance, partner across engineering and research teams, maintain software quality through testing and continuous delivery, and mentor engineers on distributed systems and backend best practices.
2 Days AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform across frontend, CLI, APIs, and backend systems. Build features using React, Python, or Go; improve stability, performance, architecture, automation, and continuous delivery. Collaborate with engineering, product, and design leaders, own end-to-end feature development, reduce technical debt, and mentor engineers on system design and problem-solving.
2 Days AgoSaved
Hybrid
4 Locations
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform’s frontend and UI infrastructure using React and Redux. Build end-to-end features, improve stability and performance, evaluate technical architecture, automate software delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a rapidly changing SaaS environment while maintaining high standards for code quality and system design.