Jobs in Bellevue, WA

833verified openings
Filter by
No filters selected
designworkstalent Verified 3h ago

AI Training Infrastructure Engineer

Bellevue, Washington, United States Hybrid

Salary not listed
PaySalary not listed
TypeFull-time
Work settingHybrid
Verified listing

JobFig found this opening at its original source and checks that it remains available.

About the role

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale. This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable. You'll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

What you'll bring

  • You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience working with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving engineering environment.
  • Comfortable operating with high ownership and limited process overhead.
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
  • Experience optimizing GPU utilization, training performance, or distributed system reliability.

Benefits

401(k)Paid HolidaysVision Insurance