Fuse Energy Logo

Fuse Energy

HPC Network Engineer

Reposted Yesterday
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Senior level
In-Office
London, Greater London, England, GBR
Senior level
Design, deploy, and operate lossless RDMA-capable leaf-spine fabrics for multi-tenant GPU clusters. Automate provisioning, build telemetry/observability, troubleshoot end-to-end performance, manage out-of-band and office networks, support tenant onboarding, and document designs while upskilling colleagues.
The summary above was generated by AI

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

The Opportunity

You'll design, deploy, and operate the network fabric for our multi-tenant AI cluster. This covers the full stack: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, fire walling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable. Beyond the data centre, you'll own the office network and act as the networking authority for the company, raising the bar for everyone by sharing what you know. You'll own the fabric from architecture through day-2 operations.

Responsibilities

  • Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale
  • Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)
  • Implement and maintain per-tenant network isolation across compute, storage, and management planes.
  • Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI
  • Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry)
  • Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts
  • Operate the out-of-band management network, console access, and remote recovery paths
  • Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales
  • Write clear design documentation capturing decisions, rationale, and rejected alternatives
  • Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments
  • Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently

Requirements
  • 5+ years as a network engineer operating production data centre networks
  • Strong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN)
  • Hands-on experience with leaf-spine / Clos fabric design and operation
  • Experience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLI
  • Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, with an understanding of why lossless behaviour matters for GPU workloads
  • Network automation as a working practice, not an aspiration: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changes
  • Solid Linux administration fundamentals: you can debug from the host side as well as the switch side
  • Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)
  • Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access (e.g. Cisco Catalyst/Meraki or comparable)
  • Clear communicator who enjoys teaching: able to document, pair, and run sessions that bring less network-savvy colleagues up to speed

Nice to have

  • Experience with GPU cluster networking specifically (e.g. NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP, or equivalent Broadcom/Arista AI fabric platforms)
  • Container networking experience: CNI plugins and BGP integration between clusters and the fabric.
  • Understanding of collective-communication libraries and how fabric behaviour shows up as training/inference performance
  • Multi-tenant network design: VRF-based isolation, tenant bandwidth guarantees, secure shared infrastructure
  • Experience with enterprise firewall platforms (e.g. FortiGate, Palo Alto), including HA deployment and virtualised/segmented instances
  • Storage networking experience (e.g. NVMe-oF, lossless storage fabrics, per-tenant storage isolation)
  • Bare-metal provisioning environments (e.g. MAAS, PXE, Redfish) and how network bootstrap fits into node lifecycle
  • Optical layer knowledge at 200/400/800G: transceivers, MPO cabling, link qualification
  • Experience standing up a data centre network from greenfield
  • Relevant certifications (e.g. CCNP/CCIE or equivalent), valued as evidence of depth, not a gate

Benefits
  • Competitive salary and an equity sign-on bonus
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Breakfast and dinner allowance for office-based employees

Similar Jobs

49 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Own Perk’s US and UK PR strategy, building Tier 1 relationships with business and financial media. Lead proactive pitching, crisis communications, corporate narrative development, executive profiling, media training, and translation of AI and technology developments into compelling stories. Manage PR agencies, measure coverage and ROI, and integrate AI tools into research, drafting, monitoring, and reporting. Partner closely with senior leadership, product, engineering, and global communications teams in a fast-paced, hybrid environment.
Top Skills: Artificial Intelligence (Ai)SaaS
57 Minutes Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain resilient systems, automation, operational APIs, observability, and infrastructure-as-code tooling. Diagnose distributed-system incidents from edge to origin, manage monitoring and alerting platforms, configure Cloudflare services, participate in incident response and post-mortems, and improve reliability, performance, and operational consistency. Collaborate across SRE, development, and IT Operations while mentoring colleagues and applying AI tools to increase productivity and root-cause analysis.
Top Skills: AnsibleCdnCloudflareCoding AssistantsDdos ProtectionDnsGoGrafanaInfrastructure As CodeJavaScriptLlm PlatformsNew RelicOpentelemetryPagerdutyPythonShell ScriptingSplunkTerraformWaf
An Hour Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Leads the re-architecture and delivery of business-critical risk and regulatory platforms using Golang, React, cloud technologies, Kafka, SQL, .NET, and TypeScript. Provides technical leadership, governance, solution design, code-quality oversight, estimation, documentation, escalation support, and mentoring. Ensures highly available, scalable, maintainable, performant, and cross-browser-compatible systems while collaborating with stakeholders and architects from development through production.
Top Skills: .NetCloud PlatformsGoKafkaReactSQLTypescript

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account