W

Senior AI Infrastructure Engineer (Coding Agent & LLM Inference Systems)

workorai • United State
Visa Sponsorship
Apply
AI Summary

Build and optimize an autonomous coding agent that develops, validates, and improves LLM inference stack components—including GPU kernels, runtime, and serving infrastructure—with strict correctness and performance requirements. Design evaluation loops, orchestrate parallel development, and ensure hardware-validated correctness. Requires deep expertise in production AI agents, LLM inference systems, or GPU kernel development with strong Python and low-level systems coding skills.

Key Highlights
Develop an autonomous agent controller for LLM inference stack components with zero-tolerance for incorrect output
Design evaluation loops for correctness validation against reference implementations and hardware benchmarks
Own test infrastructure and orchestrate parallel development across kernels, runtime, and serving infrastructure
Key Responsibilities
Build the central agent controller that writes, executes, evaluates, and improves inference code for LLM stacks
Design evaluation loops to detect incorrect output and enforce performance gates via reference comparisons and hardware benchmarks
Own the test ladder for validating correctness and performance before integrating with higher-level components
Orchestrate parallel development across GPU kernels, runtime components, and serving infrastructure with retry, dependency, and escalation logic
Review and validate low-level code produced by the agent against real hardware and simulation results
Technical Skills Required
Python LLM Inference Systems GPU Kernel Development
Benefits & Perks
$200,000–$420,000 compensation range
Visa sponsorship (including H-1B)
San Francisco preferred; remote work may be discussed for strong candidates
Nice to Have
LLVM
MLIR
TVM
Compiler code generation
Accelerator bring-up
NCCL
Megatron-LM
DeepSpeed
Distributed inference
Firmware
Non-GPU execution models
C++
Rust
CUDA C
ROCm/HIP
Metal
vLLM
KV cache
Paged attention
Continuous batching
Quantization
Speculative decoding
FlashAttention
Low-latency serving

Job Description


WorkorAI is recruiting on behalf of an early-stage AI infrastructure company building a coding agent that creates and improves an entire LLM inference stack: GPU kernels, runtime, and serving infrastructure.


This is a senior engineering role for someone whose experience combines production AI agents with either LLM inference systems or GPU kernel development.


About the role


The company is developing an agent that receives specifications and test results, writes low-level code, runs it against real hardware or simulation, analyzes failures, and keeps iterating until strict correctness and performance gates pass.


There is no room for plausible-looking output that does not work. Every component is validated against reference implementations, test suites, and hardware benchmarks.


On its first target, a proprietary accelerator with no existing inference ecosystem, the system reached working tensor-parallel matrix multiplication in approximately 10 hours and ran three frontier models end to end within 10 days. It is now serving production traffic.


What you will do


• Build the central agent controller that writes, executes, evaluates, and improves inference code.


• Design evaluation loops that detect incorrect output, compare implementations against references, and enforce performance gates.


• Own the test ladder used to validate every layer before other components are built on top of it.


• Orchestrate parallel work across kernels, runtime components, and serving infrastructure, including retry, dependency, escalation, and human-review logic.


• Review and validate low-level code produced by the agent against real hardware and simulation results.


What we are looking for


You have production experience building with LLMs or coding agents, including tool-use loops, constrained code generation, evaluation harnesses, and systems where tests catch model errors.


You also have meaningful depth in at least one of these areas:


• LLM inference systems: vLLM, KV cache, paged attention, continuous batching, quantization, speculative decoding, FlashAttention, or low-latency serving.


• GPU and accelerator kernels: CUDA, Triton, ROCm/HIP, Metal, attention, matrix multiplication, normalization, MoE, kernel fusion, or performance optimization.


Strong Python skills are required. You should also be comfortable reading or reviewing C++, Rust, CUDA C, or similarly low-level systems code.


Experience with LLVM, MLIR, TVM, compiler code generation, accelerator bring-up, NCCL, Megatron-LM, DeepSpeed, distributed inference, firmware, or non-GPU execution models is valuable but not required.


Why this role


• $200,000–$420,000 compensation range.


• San Francisco preferred; remote work can be discussed for a strong fit.


• Visa sponsorship may be available, including H-1B.


• $15M raised from investors focused on AI infrastructure and silicon.


• Approximately 14 engineers, with engineering and product operating as one team.


• Founded by a former Google Brain researcher.


• Production systems with measurable correctness and performance feedback rather than demo-only agent workflows.


How to apply


Apply through WorkorAI:


https://workorai.com/candidate/apply/cmsrca19o0001tpbqjf9vw7kl


The application includes a WorkorAI profile and a short role-focused AI interview. The complete process should take no more than 15 minutes.


WorkorAI is managing sourcing and initial technical evaluation for this search.


Similar Jobs

Explore other opportunities that match your interests

SAP Integration Engineering Team Lead

Programming
•
3h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

GlobalSource IT

United State

Senior Backend Software Engineer

Programming
•
3h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

fab2

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

fab2

United State

Subscribe our newsletter

New Things Will Always Update Regularly