T

Senior Machine Learning Infrastructure Engineer (Multimodal AI Inference)

techire.® San Francisco Bay Area
Visa Sponsorship Relocation
Apply
AI Summary

Build real-time inference infrastructure for next-gen multimodal foundation models, bridging research and production. Design scalable, low-latency systems for transformers, SSMs, and hybrid architectures while collaborating with researchers. Requires expertise in distributed systems, ML serving, and productionizing cutting-edge AI models.

Key Highlights
Build low-latency inference and serving infrastructure for frontier multimodal AI models (text, audio, video).
Design scalable, distributed systems for production AI workloads with high reliability and zero-to-one ownership.
Collaborate closely with researchers to productionize new model architectures beyond transformer limitations.
Key Responsibilities
Build real-time multimodal AI inference infrastructure capable of processing text, audio, and video streams.
Design scalable, distributed systems for serving foundation models (Transformers, SSMs, and hybrid architectures).
Develop monitoring and observability tools for the inference stack to ensure low-latency, high-reliability production AI workloads.
Work closely with researchers to transition new model architectures from research to production environments.
Shape the technical direction of the inference stack with significant ownership from day one.
Technical Skills Required
Large-scale distributed systems ML inference pipelines Production ML serving
Benefits & Perks
Competitive salary + equity
Fully covered medical, dental, and vision insurance
401(k) retirement plan
Nice to Have
Experience with vLLM
Experience with SGLang
Experience with Continuous Batching
Experience with CUDA
Experience with Triton

Job Description


Help build the inference stack behind the next generation of multimodal foundation models.


Most inference roles are about making existing models run faster.


This one is about helping define how entirely new model architectures are served at scale.


You'll be building real-time multimodal AI capable of processing enormous streams of text, audio and video. The research is pushing beyond today's transformer limitations, and your work will make those models usable in production.


You'll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models to run with low latency, high reliability and at scale. If you enjoy solving systems problems where every millisecond matters, you'll feel at home here.


Your focus

  • Build low-latency inference and serving infrastructure for foundation models across Transformers, SSMs and hybrid architectures.
  • Design scalable, reliable distributed systems that support production AI workloads.
  • Develop monitoring and observability across the inference stack.
  • Work closely with researchers to productionise new model architectures.
  • Help shape technical direction with significant ownership from day one.

You'll bring

  • Strong software engineering fundamentals and experience building large-scale distributed systems.
  • Experience with ML inference pipelines or serving generative models in production.
  • The ability to work through ambiguous technical challenges and deliver zero-to-one systems.
  • Experience implementing modern machine learning research into production environments.


Experience with vLLM, SGLang, Continuous Batching, CUDA or Triton would be particularly valuable but isn't essential.


Salary: Competitive + equity.


Location: San Francisco (onsite).


Alongside compensation, you'll receive fully covered medical, dental and vision insurance, 401(k), relocation and immigration support, plus daily meals in the office.


If you're interested in building the infrastructure that enables the next generation of AI models to run in real time, we'd love to tell you more.


All applicants will receive a response.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

poseidon aerospace

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Executive

confidential company

San Francisco Bay Area

Senior Backend Engineer

Programming
1d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

retell

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly