R

Senior Infrastructure Engineer

reactor United State
Visa Sponsorship Relocation
Apply
AI Summary

Own and operate the multi-cloud Kubernetes infrastructure supporting AI model workloads, including GPU orchestration and real-time networking. Lead the expansion of infrastructure across AWS and GPU cloud providers using GitOps and infrastructure-as-code. Requires deep production experience with Kubernetes, GPU scheduling, and observability stacks.

Key Highlights
Own end-to-end infrastructure for AI models across multi-region Kubernetes clusters and GPU clouds
Lead expansion to new cloud providers and manage GPU node infrastructure, scheduling, and observability
Operate complex networking layers including media relays, load balancing, and cross-region connectivity
Key Responsibilities
Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code
Own the GitOps deployment lifecycle including Helm charts, Kustomize overlays, image automation, and continuous delivery
Manage GPU node infrastructure including scheduling, model weight caching, image prefetching, and GPU observability
Operate and improve the networking layer including ingress, gateway management, load balancing, media relay infrastructure, and cross-region connectivity
Build and maintain the observability stack for metrics, logs, traces, and profiling across services and GPU workloads
Maintain infrastructure security including IAM, secret management, certificate automation, and encryption at rest
Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases
Partner with ML engineers on model serving including container optimization, health checks, media pipeline performance, and multi-GPU configuration
Technical Skills Required
Kubernetes Infrastructure-as-Code GPU Workloads
Benefits & Perks
Competitive San Francisco salary and meaningful equity
Visa sponsorship and US relocation support
Generous health, dental, and vision coverage
Nice to Have
Experience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)
Real-time media or streaming infrastructure experience
Go or Python proficiency
Familiarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)
FinOps and GPU cost optimization experience

Job Description


Department: Engineering

Location: San Francisco

Description

You'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets.

We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running.

What You'll Do

  • Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code.
  • Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.)
  • Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability.
  • Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity.
  • Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads.
  • Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest.
  • Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases.
  • Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration.

What We're Looking For

  • You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades.
  • Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions.
  • You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing matters
    for expensive hardware.
  • Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring.
  • Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns)
  • Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces.
  • Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus.
  • You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it.
Nice to Have
  • Experience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)
  • Real-time media or streaming infrastructure
  • Go or Python proficiency
  • Familiarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)
  • FinOps and GPU cost optimization
What We're Not Looking For
  • Pure CI/CD pipeline engineers who haven't operated Kubernetes clusters directly
  • Candidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway)
  • People who need a fully defined scope, this role requires figuring out what to build next, not just executing tickets

Benefits

  • Competitive San Francisco salary and meaningful equity
  • We sponsor visas and support relocation to the US
  • Generous health, dental, and vision coverage

Similar Jobs

Explore other opportunities that match your interests

Resiliency Engineer

Devops
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

vanguard

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Palo Alto Networks

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

inventure

United State

Subscribe our newsletter

New Things Will Always Update Regularly