Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
US, TX, AustinAll locationsUS, TX, AustinUS, CA, Santa ClaraUS, OR, RemoteUS, CA, Remote Remote
Salary not listedJobFig found this opening at its original source and checks that it remains available.
About the role
We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal point for end-to-end performance, working directly with engineering teams of key CSP/hyperscale customers to ensure they achieve various performance targets on NVIDIA platforms. In this role, you will augment NVIDIA's performance and benchmark teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of platform performance characteristics, gather and incorporate their workload-specific feedback into NVIDIA's optimization priorities, and validate that performance targets are met in customer-representative configurations.
What you'll bring
- 15+ years of experience in systems performance engineering
- BS or MS in Computer Science, Computer Engineering, or related field (or equivalent experience)
- Proficiency in GPU workload profiling: nsight systems, nsight compute, DCGM metrics, or equivalent instrumentation
- Understanding of distributed training performance dynamics: computation/communication overlap, pipeline bubbles, memory bandwidth utilization, collective efficiency
- Statistical methods for performance analysis: regression detection, confidence intervals, A/B comparison at scale
- Understanding of how the full software stack impacts performance: driver overhead, collective algorithm selection, memory allocation, scheduling, firmware power management
- Strong data analysis and visualization skills (Python, pandas, dashboards).
- Customer obsession — genuine passion for understanding why customers aren't achieving expected performance and driving solutions
- Ability to communicate performance findings to both deep technical audiences and executive leadership
- Demonstrated success influencing multiple engineering teams to prioritize performance improvements
- ideally in GPU/HPC/ML infrastructure.