deeter-analytics Verified 30h ago
Head of AI Inference & MLOps
Austin, Texas, United States Hybrid
Salary not listedPaySalary not listed
TypeFull-time
Work settingHybrid
Verified listing
JobFig found this opening at its original source and checks that it remains available.
About the role
This is not a traditional datacenter operations role. We are hiring the person who will make the racks make money.
What you'll bring
- Significant experience in production AI/LLM inference, MLOps, model serving, or AI infrastructure monetization
- Proven experience running or scaling GPU-backed inference systems in production
- Strong understanding of modern inference runtimes, serving frameworks, and optimization techniques
- vLLM
- TensorRT-LLM
- SGLang
- Ray Serve
- Triton Inference Server
- Kubernetes-based GPU orchestration
- custom routing / scheduler layers
- Experience optimizing for real-world production metrics such as throughput, latency, GPU utilization, availability, and cost efficiency
- Strong understanding of LLM inference economics, including tradeoffs among model size, quantization, latency, throughput, memory footprint, and customer willingness to pay