S

Member of Technical Staff - AI Inference Engineer

Simplify • United State
Visa Sponsorship
Apply
AI Summary

Develop and optimize high-performance computing kernels for LLM inference systems. Own customer accounts end-to-end from technical evaluation to production operations. Must be in-person in San Francisco 5 days/week.

Key Highlights
Own 3-4 named customer accounts directly
Build and operate heterogeneous clusters across vendors
Win technical evaluations on customer workloads
Key Responsibilities
Develop and optimize high-performance computing kernels for inference engine internals and serving infrastructure
Design, deploy, and operate heterogeneous clusters across vendors
Own customer accounts end-to-end from technical evaluation to production operations
Win technical evaluations on customer workloads
Inform product roadmap based on real-world system behavior
Technical Skills Required
Python C++ Linux AI inference
Benefits & Perks
$200-300K base salary + equity
Housing stipend ($1K/month)
Covered Uber/Waymo rides
Nice to Have
vLLM
SGLang
TensorRT-LLM
Triton
Early-stage or founding experience

Job Description


About the role

Most job posts are designed to collect as many applications as possible.


This is not one of them.


At Simplify, we work directly with startup founders to help them hire for high-priority roles.


We're partnering with a fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference fast, and own the customers running on them. $200–300K base + equity, and they're hiring multiple engineers now.


About the company

One-liner: Building the infrastructure that makes LLM inference fast — and putting the engineers who build it directly in front of the customers who run on it


Stage: Small team, HQ in SF. Their customers include Vercel and Neon Health, and the engineers own those relationships directly — no solutions team, no layer between you and the workload.


Most infra engineers ship into a backlog and never meet the people running their code. Here, sales gets the first meeting — from there, you decide what to prove, build it, win the technical evaluation on the customer's own workload, and keep it running in production. Each engineer owns 3–4 named accounts outright.


What you'll work on

  • Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure
  • Develop AI agents to do autonomous inference engineering
  • Design, deploy, and operate heterogeneous clusters across vendors
  • Own customer accounts end to end. Sales gets the first meeting. From there you decide what to prove, you build it, you keep it running in production, and you carry the relationship
  • Win the technical evaluation. Prove Wafer on the customer's own workload rather than a synthetic benchmark, and be the person who can explain the result to their engineers
  • Own production for your accounts. When latency moves or error rates climb, you find it, you fix it or route it, and you are who the customer hears from
  • Inform product roadmap. You sit closer than anyone to how Wafer behaves under real load, and the engineering team builds against what you report


What we look for

  • You've shipped and operated production systems — been on call, personally debugged production problems, and can walk through how you found them
  • You've owned a customer relationship or a deployment with real users — you can name it, and you were the person people came to
  • Credible in a room full of engineers — you can defend a benchmark methodology
  • Able to work without a spec
  • In person in San Francisco, 5 days/week
  • Already in the US (they sponsor and transfer essentially any visa status — H-1B, O-1, TN, STEM OPT)


Nice to have

  • Inference / model serving experience (vLLM, SGLang, TensorRT-LLM, Triton) — their most useful differentiator, but they'll interview a strong engineer without it
  • Early-stage or founding experience — if you've sold your own product to engineers, you've already done both halves of this job
  • Heavy use of AI agents and tooling in your own workflow


Why join

  • $200–300K base + equity — plus a $1K/month housing stipend near the office and covered Uber/Waymo rides
  • No years-of-experience requirement in either direction — the bar is what you've operated, not how long you've worked
  • Your work shows up directly in customer production systems, and you carry the relationship
  • Small fast-moving team, high ownership — you own named accounts outright from day one
  • They're moving fast: offers signed by end of August



Similar Jobs

Explore other opportunities that match your interests

Robot Application Developer

Programming
•
39m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

Senior ML Systems Engineer (Reinforcement Learning & LLM Finetuning)

Programming
•
52m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

anthropic

United State

Founding Engineer - Full Stack

Programming
•
57m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

speedyapply

United State

Subscribe our newsletter

New Things Will Always Update Regularly