Lead the design, deployment, and operation of Callosum's heterogeneous AI compute clusters across multiple sites. Scope and commission new infrastructure while building and managing a team for reliability and observability. Requires deep expertise in power, cooling, rack design, networking, and storage integration.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
About Us
We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.
Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.
The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.
Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.
Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.
In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.
We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.
About The Role
Looking to advance your IT & Network Engineering career with relocation support? Explore IT & Network Engineering Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
This role owns Callosum’s physical compute infrastructure, from the first assessment of a potential deployment through to the reliable operation of the resulting cluster. You will scope sites, coordinate the design cluster infrastructure and deployment across vendors and facilities, and lead commissioning of new hardware. As our footprint grows, you will build and lead the team responsible for keeping that infrastructure healthy, observable, and available, establishing the operational systems that turn individual deployments into a reliable fleet.
What You'll Do
- Scope new cluster deployments across site selection, power, cooling, rack layout, networking, storage, fibre connectivity, and capacity requirements
- Coordinate delivery and commissioning across facilities, utilities, OEMs, networking and storage vendors, ensuring the physical and technical dependencies come together
- Build the operational infrastructure for the fleet, including hardware telemetry, health monitoring, maintenance, incident response, capacity planning, and hardware lifecycle management
- Build and lead the physical infrastructure team responsible for deploying new clusters and maintaining the reliability and availability of the resources we operate
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Experience designing, deploying, or operating high-density compute infrastructure, HPC clusters, AI infrastructure, or similarly complex physical systems
- Strong understanding of how power, cooling, rack design, networking, storage, and compute interact to determine the capabilities and constraints of a cluster
- Demonstrated ability to drive complex infrastructure deployments across technical teams, facilities, hardware vendors, network providers, and other external stakeholders
- Experience building and leading infrastructure teams, with strong operational instincts around reliability, observability, capacity, maintenance, and failure management
Interested in relocating to United Kingdom? Check out our comprehensive Relocation Jobs in United Kingdom page with detailed relocation packages and benefits.
- Competitive Salary, determined by skills and experience
- Equity & Ownership
- Private healthcare
- We offer Visa sponsorship and relocation benefits to hire the best in the world
- We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us
Similar Jobs
Explore other opportunities that match your interests