C

Senior Systems Communication Engineer (Heterogeneous AI Accelerators)

callosum United Kingdom
Visa Sponsorship Relocation
Apply
AI Summary

Lead the development of high-performance communication systems for heterogeneous AI accelerators, optimizing data movement across protocols, hardware, and software layers to eliminate latency bottlenecks. Own technical strategy for RDMA, programmable NICs, and emerging interconnects to enable scalable, cost-efficient AI inference. Drive innovation in cross-device memory management and cluster architecture for next-gen intelligent systems.

Key Highlights
Own technical strategy for reducing communication overhead in heterogeneous AI systems, spanning protocols, memory, and networking hardware
Develop high-speed data movement solutions using RDMA, programmable NICs, DPUs, and cross-device memory management
Evaluate and integrate emerging scale-up/scale-out fabrics to inform cluster architecture and accelerator communication evolution
Key Responsibilities
Characterize and simplify communication paths between heterogeneous accelerators, tracing latency and overhead across networking, runtime, memory, and device layers
Develop faster data movement mechanisms across accelerators using RDMA, programmable NICs, DPUs, and software-based cross-device memory management
Build flexible model state transfer mechanisms (e.g., KV cache connectors) for different accelerator types and execution strategies
Evaluate existing and emerging scale-up/scale-out fabrics to inform cluster architecture and heterogeneous accelerator communication strategies
Technical Skills Required
Networking and Communication Systems High-Performance Data Movement (RDMA, Ethernet fabrics) Systems Performance Optimization
Benefits & Perks
Competitive Salary (£101K - £192K)
Equity & Ownership
Private healthcare

Job Description


About Us

We’re living through a Cambrian explosion of intelligence: new models and new chips, each specialised for different tasks, are arriving all at once. The result is a new era for AI, one of radical heterogeneity.

Callosum is the Intelligent Systems Company. We believe the next generation of AI won't be defined by any single model or chip, but by intelligent systems in which hardware and intelligence co-evolve. We are building the infrastructure that unifies heterogeneous compute across the full stack. This opens a new axis of scaling intelligence: a dynamic system that tailors itself to what each workload actually needs, whether that's speed, cost, precision, or whatever unit comes next.

The last era scaled on a different bet: one bigger model, more of the same chip, more data. That bet is running into structural limits. Frontier models offer extraordinary capability at unsustainable cost, one that today's monolithic infrastructure was never designed to serve.

Our founding principle is that intelligence comes from many specialised systems working together, not from any single component. We build the software orchestration layer that co-evolves models, workflows and silicon into one system, delivering inference tailored to every workload, and demonstrating orders-of-magnitude leaps in capability and cost.

Because our software spans the full stack, our engineering team works directly with heterogeneous accelerators and frontier silicon, including Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, and Mixx. We are not stopping at today's chips: each new generation of silicon unlocks algorithms that couldn't run before, and we intend to be first to them, every time. If we get it right, it will belong to everyone building on it - not to any single vendor.

In our latest funding round, we raised $100M, led by Atomico with participation from Plural, DCVC and the UK Sovereign AI Fund’s first investment. With this, we are building the infrastructure for the next era of intelligence.

We are engineers and scientists based in London, working across the full depth of the stack. We are curious, intellectually honest, and building what doesn't exist yet. If you thrive on uncharted territory and are energised by the scale of the challenge, we'd love to hear from you.

About The Role

Callosum believes that orders of magnitude improvements in AI systems will come through application-aware orchestration across heterogeneous hardware. As workloads become increasingly disaggregated, the performance of the system is determined not only by where computation happens, but by how efficiently state and data can move between the hardware best suited to each part of the workload. Communication latency sets the boundary on how far that disaggregation can go.

This role owns that boundary. You will develop the communication systems that allow heterogeneous accelerators to operate as parts of one system, working across protocols, memory movement, networking hardware, and the software layers above them. That means understanding where overhead actually comes from, removing assumptions inherited from more general-purpose workloads, and finding new ways to move data directly and efficiently across vendor boundaries. The remit is intentionally broad: from RDMA paths and programmable NICs to cache transfer formats and emerging interconnect technologies, you will own the technical strategy for making communication less of a constraint on the systems we can build.

What You'll Build

  • Characterise and simplify communication paths between accelerators, tracing latency and unnecessary overhead across networking, runtime, memory, and device layers
  • Develop faster data movement across heterogeneous hardware, exploring RDMA, programmable NICs, DPUs, switches, and software approaches to cross-device memory management
  • Build flexible mechanisms for transferring model state across accelerator types and execution strategies, including KV cache connectors across different forms of parallelism
  • Evaluate existing and emerging scale-up and scale-out fabrics, informing cluster architecture, infrastructure integration, and how heterogeneous accelerator communication should evolve

What Sets You Apart

  • Deep understanding of networking and communication systems, with the ability to reason across protocols, hardware, memory systems, and software layers
  • Experience with high-performance data movement such as RDMA, Ethernet fabrics, accelerator interconnects, programmable networking, or similar latency-sensitive systems
  • Strong systems performance instincts, with the ability to trace bottlenecks across abstraction layers and distinguish fundamental constraints from accidental ones
  • A demonstrable tendency to question existing abstractions, generalise across unfamiliar technologies, and simplify systems around the requirements of the workload

What We Offer

  • Competitive Salary, determined by skills and experience
  • Equity & Ownership
  • Private healthcare
  • We offer Visa sponsorship and relocation benefits to hire the best in the world
  • We work in person at our London office. You'll have the tools, space and setup to do your best work, and if you have specific needs, just tell us

We're committed to building an inclusive workplace where everyone feels welcome, and believe in equal opportunities for all.

Compensation Range: £101K - £192K


Similar Jobs

Explore other opportunities that match your interests

Storage Systems Engineer

Networking
11h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

European Bioinformatics Instit...

United Kingdom
Visa Sponsorship Relocation Remote
Job Type Internship
Experience Level Entry level

durham university

United Kingdom

Controls and Instrumentation Engineer

Networking
6d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Lonza

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly