Top Tech Jobs & Startup Jobs in London

2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Design and build backend systems for AI agent orchestration, tool execution, workflow management, memory, and state handling. Develop scalable APIs and reliable infrastructure for distributed AI workflows, improve platform reliability and observability, and partner with product, research, and engineering teams to deliver agent capabilities. Own technical initiatives from architecture through production, maintain software quality through testing and continuous delivery, and mentor engineers on backend and distributed-systems practices.
Top Skills: APIsAsynchronous ProcessingAutomated TestingAWSAzureCloud-Native ArchitectureContinuous DeliveryDistributed SystemsDockerGCPGoKubernetesObservabilityPythonRust
2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software
Develop and scale Lightning AI’s backend platform using Go, spanning APIs, infrastructure, billing, security, and integrations. Own features end to end, design reliable and scalable systems, improve architecture and performance, maintain software quality and continuous delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a fast-changing SaaS environment.
Top Skills: AWSAzureDockerGCPGoKubernetes
2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software
Support ML engineering teams running production training and inference workloads across Kubernetes, cloud, Linux, and GPU infrastructure. Diagnose distributed PyTorch, CUDA, NCCL, networking, storage, scheduling, and performance issues; analyze observability data; advise customers during incidents; and improve reliability through tooling, automation, documentation, runbooks, and operational processes. The role is hybrid in London, requires at least two office days weekly, and supports EMEA shifts.
Top Skills: Bare Metal InfrastructureCloud InfrastructureContainerizationCudaDistributed SystemsGpu InfrastructureGrafanaInfinibandKubeflowKubernetesLinuxNcclOpentelemetryPrometheusPythonPyTorchRayRdmaSlurm
2 Days AgoSaved
Hybrid
London, Greater London, England, GBR
Mid level
Mid level
Artificial Intelligence • Machine Learning • Software
Build and deploy production AI systems for customers, translating business objectives into scalable technical solutions. Responsibilities include architecture, proof-of-concepts, software development, deployment, monitoring, debugging, inference optimization, and distributed systems operation. The engineer partners with customer engineering teams, collaborates with product and engineering, improves reusable platform capabilities, and owns technical engagements from discovery through production scaling.
Top Skills: APIsDistributed SystemsDockerGoGpu-Accelerated WorkloadsKubernetesLanggraphModel Serving SystemsPythonRayReactTensorrtTypescriptVector DatabasesVllmWorkflow Orchestration Systems
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
Top Skills: APIsBare-Metal InfrastructureBmcContainerizationDell HardwareGpu ServersHpcIpmiJuniper NetworksLinuxOrchestrationPalo Alto FirewallsPxe/IpxePythonRedfishSonicVast
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software
This is a general talent community opportunity rather than a specific open position. Lightning AI invites candidates to submit their information for consideration as future roles become available across U.S. and London hubs, with occasional remote opportunities. The company develops tools and infrastructure for building, training, and deploying AI systems and values ownership, urgency, communication, teamwork, continuous improvement, and long-term thinking.
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Operate, scale, and optimize distributed storage infrastructure supporting large-scale AI/ML and HPC workloads. Build Python automation, manage Linux bare-metal systems, troubleshoot storage, hardware, networking, and operating system issues, and improve performance, reliability, monitoring, capacity planning, and lifecycle management. Collaborate with infrastructure, networking, platform, and data center teams on storage deployments and scaling strategies.
Top Skills: CephGpu Direct StorageLinuxNfsPythonRdmaS3Vast
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Build, validate, and operate large-scale bare-metal GPU infrastructure for AI/ML and HPC workloads. Responsibilities include managing Linux systems, image pipelines, test clusters, provisioning, firmware and driver validation, GPU diagnostics, performance analysis with NVIDIA DCGM, automation, virtualization, and hardware management interfaces. The role requires troubleshooting across hardware and software layers while collaborating with infrastructure, hardware, data center, platform, and ML teams.
Top Skills: Bare-Metal ProvisioningGpusIdracImage-Based ProvisioningInfinibandIpmiLinuxLivecdNvidia DcgmNvlinkPxePythonRedfishVirtualization
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
Top Skills: AnsibleAWSBashCephDellElk StackEthernetGitopsGoGpuInfinibandJuniperKubernetesLinuxNfsPalo AltoPrometheusPythonSonicTerraformUbuntuVast
2 Days AgoSaved
Remote or Hybrid
London, Greater London, England, GBR
Entry level
Entry level
Artificial Intelligence • Machine Learning • Software
Develop and post-train deep learning models while building systems, tooling, and workflows for training, evaluation, debugging, and deployment. Contribute to open-source projects, distributed AI infrastructure, backend services, and developer platforms. Collaborate with research, product, infrastructure, and customer teams to solve complex technical problems, prototype ideas, and productionize successful experiments.
Top Skills: Cloud InfrastructureCudaDeepspeedDistributed SystemsFsdpHugging FaceLightning AiNvidia MoltPyTorchSglangTritonVllm
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account