Lead the development of high-performance communication systems for heterogeneous AI accelerators, optimizing data movement across protocols, hardware, and software layers to eliminate latency bottlenecks. Own technical strategy for RDMA, programmable NICs, and emerging interconnects to enable scalable, cost-efficient AI inference. Drive innovation in cross-device memory management and cluster architecture for next-gen intelligent systems.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
About Us
We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.
Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.
The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.
Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.
Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.
In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.
We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.
About The Role
Looking to advance your IT & Network Engineering career with relocation support? Explore IT & Network Engineering Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
This role owns that boundary. You will develop the communication systems that allow heterogeneous accelerators to operate as parts of one system, working across protocols, memory movement, networking hardware, and the software layers above them. That means understanding where overhead actually comes from, removing assumptions inherited from more general-purpose workloads, and finding new ways to move data directly and efficiently across vendor boundaries. The remit is intentionally broad: from RDMA paths and programmable NICs to cache transfer formats and emerging interconnect technologies, you will own the technical strategy for making communication less of a constraint on the systems we can build.
What You'll Build
- Characterise and simplify communication paths between accelerators, tracing latency and unnecessary overhead across networking, runtime, memory, and device layers
- Develop faster data movement across heterogeneous hardware, exploring RDMA, programmable NICs, DPUs, switches, and software approaches to cross-device memory management
- Build flexible mechanisms for transferring model state across accelerator types and execution strategies, including KV cache connectors across different forms of parallelism
- Evaluate existing and emerging scale-up and scale-out fabrics, informing cluster architecture, infrastructure integration, and how heterogeneous accelerator communication should evolve
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Deep understanding of networking and communication systems, with the ability to reason across protocols, hardware, memory systems, and software layers
- Experience with high-performance data movement such as RDMA, Ethernet fabrics, accelerator interconnects, programmable networking, or similar latency-sensitive systems
- Strong systems performance instincts, with the ability to trace bottlenecks across abstraction layers and distinguish fundamental constraints from accidental ones
- A demonstrable tendency to question existing abstractions, generalise across unfamiliar technologies, and simplify systems around the requirements of the workload
Interested in relocating to United Kingdom? Check out our comprehensive Relocation Jobs in United Kingdom page with detailed relocation packages and benefits.
- Competitive Salary, determined by skills and experience
- Equity & Ownership
- Private healthcare
- We offer Visa sponsorship and relocation benefits to hire the best in the world
- We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us
Compensation Range: £101K - £192K
Similar Jobs
Explore other opportunities that match your interests