Staff Software Engineer - Managed Kubernetes
San Jose, California, United StatesAll locationsSan Jose, California, United StatesBellevue, Washington, United StatesSan Francisco, California, United States Hybrid
$314,000–$419,000 a yearJobFig found this opening at its original source and checks that it remains available.
About the role
We are seeking a Staff Engineer to help our development of our Managed Kubernetes platform. Think GKE, but purpose-built for AI workloads and running on bare metal. This is a foundational technical leadership role where you will shape the infrastructure that powers the next generation of AI training and inference at scale. As a Staff Engineer on our Orchestration team, you will collaborate to help drive the technical vision for Lambda's managed orchestration services, including Managed Kubernetes, Managed Slurm on Kubernetes, and higher-level platform services for inference and AIOps. You'll work at the intersection of distributed systems, GPU-accelerated computing, and Cloud Native infrastructure to build systems that are reliable, performant, and elegantly simple for our customers. This is not a role for someone who just operates Kubernetes; it is a technical leadership role for an engineer who has synthesized the core domains of infrastructure (compute, network, storage, security) and can design holistic solutions across all of them. You'll be working closely with NVIDIA's open-source ecosystem, and partnering with internal teams across the stack to deliver a world-class managed platform. We are seeking a Staff Engineer to help our development of our Managed Kubernetes platform. This is a foundational technical leadership role where you will shape the infrastructure that powers the next generation of AI training and inference at scale.
What you'll bring
- You are a creative, innovative engineer who operates at high velocity.
- You don't just solve problems.
- You find elegant solutions and ship them quickly.
- You embrace modern tools and AI-assisted development (like Claude Code) to accelerate your productivity and multiply your impact.
- You're energized by building new things, not maintaining the status quo.
- 10+ years of experience in software engineering, platform engineering, or SRE, with at least 5 years focused on Kubernetes at scale
- Expert-level understanding of Kubernetes internals: API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
- Holistic infrastructure expertise: you've synthesized knowledge across compute, networking, storage, and security, not just Kubernetes in isolation.
- You can build solutions that span the full stack.
- Strong software engineering skills in Go (required) and Python
- you write production-quality code, not just scripts
- Deep experience with GPU orchestration in Kubernetes: NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.