Callosum Logo

Callosum

Head of Physical Infrastructure

Posted 8 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Entry level
In-Office
London, Greater London, England, GBR
Entry level
Owns Callosum’s physical compute infrastructure from site assessment and cluster design through deployment, commissioning, and reliable operation. Responsibilities include coordinating power, cooling, rack, networking, storage, fiber, facilities, utilities, and vendor dependencies; building fleet-wide telemetry, monitoring, maintenance, incident response, capacity planning, and hardware lifecycle systems; and hiring and leading the physical infrastructure team.
The summary above was generated by AI

About Us

We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.

Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.

The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.

Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.

Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.

In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.

We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.

About the Role

Callosum believes that the next generation of AI infrastructure will be built from a much more diverse set of hardware than the clusters of today. Scaling that infrastructure creates a new physical systems problem: different accelerators bring different power densities, cooling requirements, rack constraints, network topologies, vendor dependencies, and deployment models. We want to find ways to turn those increasingly complex inputs into physical infrastructure that can be deployed repeatedly and scaled.

This role owns Callosum’s physical compute infrastructure, from the first assessment of a potential deployment through to the reliable operation of the resulting cluster. You will scope sites, coordinate the design cluster infrastructure and deployment across vendors and facilities, and lead commissioning of new hardware. As our footprint grows, you will build and lead the team responsible for keeping that infrastructure healthy, observable, and available, establishing the operational systems that turn individual deployments into a reliable fleet.

What You'll Do

  • Scope new cluster deployments across site selection, power, cooling, rack layout, networking, storage, fibre connectivity, and capacity requirements

  • Coordinate delivery and commissioning across facilities, utilities, OEMs, networking and storage vendors, ensuring the physical and technical dependencies come together

  • Build the operational infrastructure for the fleet, including hardware telemetry, health monitoring, maintenance, incident response, capacity planning, and hardware lifecycle management

  • Build and lead the physical infrastructure team responsible for deploying new clusters and maintaining the reliability and availability of the resources we operate

What Sets You Apart

  • Experience designing, deploying, or operating high-density compute infrastructure, HPC clusters, AI infrastructure, or similarly complex physical systems

  • Strong understanding of how power, cooling, rack design, networking, storage, and compute interact to determine the capabilities and constraints of a cluster

  • Demonstrated ability to drive complex infrastructure deployments across technical teams, facilities, hardware vendors, network providers, and other external stakeholders

  • Experience building and leading infrastructure teams, with strong operational instincts around reliability, observability, capacity, maintenance, and failure management

What We Offer

  • Competitive Salary, determined by skills and experience

  • Equity & Ownership

  • Private healthcare

  • We offer Visa sponsorship and relocation benefits to hire the best in the world

  • We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us

We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.

Similar Jobs

15 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Own Expedia's end-to-end loyalty messaging strategy across discovery, shopping, booking, and post-trip engagement. Define value-led messaging hierarchies, optimize member communications, and align loyalty experiences across hotels, flights, packages, cars, and activities. Partner with commercial, design, research, analytics, engineering, and loyalty teams to create scalable frameworks. Use experimentation, measurement, and safely integrated AI/ML solutions to improve comprehension, activation, personalization, and repeat bookings.
Top Skills: Ai/Ml
15 Minutes Ago
Hybrid
London, Greater London, England, GBR
Mid level
Mid level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Monitors global crises and threat intelligence, assesses severity and traveler impact, and coordinates rapid response actions across cross-functional teams. The analyst produces intelligence briefs, exposure analyses, dashboards, metrics, and executive reports; manages incidents in ServiceNow; supports traveler communications; maintains response playbooks; conducts post-event reviews; and uses SQL, operational data, and AI tools to improve signal detection, escalation decisions, and crisis workflows.
Top Skills: Artificial Intelligence ToolsConfluenceExcelJIRALarge Language ModelsMicrosoft 365Microsoft TeamsOutlookQuerybookSalesforceServicenowSharepointSQLTrino
16 Minutes Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Design, build, and operate scalable, high-performance backend services and APIs. Lead architecture, code reviews, and cross-team integrations. Integrate AI/ML responsibly, diagnose and resolve production issues, and participate in on-call rotations to maintain operational excellence.
Top Skills: Ai-Assisted Development ToolsAWSInfrastructure As CodeJavaJvmKotlin

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account