Jobs in Palo Alto, CA

1,271verified openings
Filter by
No filters selected
SpaceXAI Verified 8h ago

Software Engineer - Training/Inference (C++)

Palo Alto, California On-site

Salary not listed
PaySalary not listed
TypeNot specified
Work settingOn-site
Verified listing

JobFig found this opening at its original source and checks that it remains available.

About the role

As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale

What you'll bring

  • Deep low-level systems programming (C/C++ or Rust)
  • Experience with large-scale, high-concurrent production serving.
  • Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
  • Strong background in system optimizations: batching, caching, load balancing, parallelism.
  • Low-level inference optimizations: GPU kernels, code generation.
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
  • Experience with testing, benchmarking, and reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.

Benefits

$180,000 - $440,000 USD