M

Research Engineer - Inference Systems

MBN Solutions United State
Visa Sponsorship Relocation
Apply
AI Summary

Build high-throughput and low-latency inference systems for novel physical-world foundation models. Optimize GPU utilization, distributed orchestration, and evaluation pipelines for large-scale backtesting. Requires deep expertise in PyTorch/JAX, GPU parallelism, and distributed computing frameworks.

Key Highlights
Work on a new class of foundation models learning from physical systems rather than text
Own the problem of making inference fast and cheap to enable rapid research evaluation
Small team (<15 people) with high funding and rapid scaling, requiring close collaboration with researchers
Key Responsibilities
Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations
Design and implement techniques for low-latency real-time inference
Optimize hardware utilization of FLOPs, bandwidth, and memory
Extend frameworks like Kubernetes, Ray, or Slurm for distributed inference and large-batch evaluation sweeps
Establish standards for reliability, observability, and reproducibility in evaluation
Collaborate with researchers to optimize novel architectures before reference implementations exist
Technical Skills Required
PyTorch JAX Distributed Computing
Benefits & Perks
Relocation support
Visa transfer support
Nice to Have
Open-source contributions to inference or systems infrastructure (vLLM, SGLang, Triton)
Distributed training with PyTorch or FSDP
Background in computer vision, robotics, sensor fusion, physics-informed ML, scientific AI, or multimodal learning

Job Description


Research Engineer - Inference | Foundation Models for the Physical World

San Francisco | On-site, five days a week | Relocation support available


Make inference so fast and cheap that evaluation never gates research.


A team of fewer than 15 people is building a new class of foundation model - one that learns cause and effect from physical systems rather than text. Models are trained from scratch, on novel architectures, across hundreds of GPUs, against petabyte-scale multimodal data drawn from one of the largest collections of real-world observational data available.

Every research decision they make depends on how quickly and cheaply those models can be evaluated. That's the problem you own.


Why this is a different problem

Most inference work today follows a well-mapped road: known architectures, known kernels, a decade of published tricks.

This isn't that. The architectures are new, the modalities are physical rather than textual, and the workload - scoring against historical observations, backtesting, large-batch evaluation sweeps - looks nothing like serving a chat endpoint. Much of what you build won't have a published answer to copy from.


If that reads as an opportunity rather than a risk, keep going.


What you'll own

  • Throughput at evaluation scale. High-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations.
  • Latency for real-time inference. Designing and implementing the techniques that make it fast enough to be useful.
  • The hardware. Real utilisation of FLOPs, bandwidth, and memory - not just the numbers the framework reports.
  • Distributed orchestration. Extending frameworks like Kubernetes, Ray, or Slurm for distributed inference and large-batch evaluation sweeps.
  • Trustworthy evaluation. Standards for reliability, observability, and reproducibility, so a number from a sweep means the same thing twice.
  • Novel architectures. Working directly with researchers to make new architectures fast, often before there's a reference implementation to work from.


What we're looking for

  • You've built or optimised inference and serving systems for throughput and latency — TensorRT, vLLM, SGLang, or something you wrote yourself.
  • You're fluent in distributed compute and GPU parallelism, and you optimise against the hardware rather than around it.
  • You know PyTorch or JAX deeply enough to reason about what they're doing underneath, not just call them.
  • You can profile an unfamiliar system, find the real bottleneck, and ship a fix that survives contact with other people's code.
  • You're comfortable spanning research and engineering, and you don't need a clean handoff to make progress.


Useful, not essential: open-source contributions to inference or systems infrastructure (vLLM, SGLang, Triton). Distributed training with PyTorch or FSDP. Background in computer vision, robotics, sensor fusion, physics-informed ML, scientific AI, or multimodal learning.

We don't expect all of it. We do expect you to learn the rest quickly.


The team

This is not a large corporate research lab. It's a highly funded early-stage company, fewer than 15 people today, scaling rapidly over the next year.

Everyone writes code. Everyone contributes to research. There's no separate infrastructure org to hand things to and no one to translate the research for you - you'll be in the room where it's decided.

Five days a week in San Francisco. That's deliberate, and it isn't negotiable - the work is too tightly coupled for anything else at this stage. Relocation and visa transfer support is available.


Similar Jobs

Explore other opportunities that match your interests

Senior Director of AI Engineering

Programming
3m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

coffeespace

United State

Senior Backend Engineer, Observability Platform

Programming
13m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State

Subscribe our newsletter

New Things Will Always Update Regularly