SpaceXAI Verified 8h ago
Software Engineer - Training/Inference (C++)
Palo Alto, California On-site
Salary not listedPaySalary not listed
TypeNot specified
Work settingOn-site
Verified listing
JobFig found this opening at its original source and checks that it remains available.
About the role
As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale
What you'll bring
- Deep low-level systems programming (C/C++ or Rust)
- Experience with large-scale, high-concurrent production serving.
- Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
- Strong background in system optimizations: batching, caching, load balancing, parallelism.
- Low-level inference optimizations: GPU kernels, code generation.
- Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
- Experience with testing, benchmarking, and reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Benefits
$180,000 - $440,000 USD