It & cybersecurity jobs in California

141verified openings
Filter by
No filters selected
sciforium Verified 38h ago

GPU Cluster Engineer, Networking

San Francisco, California, United States Hybrid

$150,000–$180,000 a year
Pay$150,000–$180,000
TypeFull-time
Work settingHybrid
Verified listing

JobFig found this opening at its original source and checks that it remains available.

About the role

We are looking for a Senior Network Engineer to own the complete networking stack of our GPU clusters — from the RDMA compute fabric to the data center perimeter to cross-site and cloud connectivity. You will design the network architecture for new large-scale clusters, lead its bring-up, and operate it in production. This role owns everything from the NIC port outward: backend InfiniBand/RoCE fabrics, frontend and storage networks, out-of-band management, firewalls and edge security, and the links that connect our clusters to each other and to the cloud. You will work alongside our Hardware Operations team (physical install) and Systems & Platform team (host software stack) to deliver a fabric that never becomes the bottleneck.

What you'll bring

  • 7+ years designing and operating production data center networks, including at least one large-scale HPC/AI cluster fabric (InfiniBand or RoCE v2) you designed or ran.
  • Expert-level routing and switching: BGP (including EVPN-VXLAN), OSPF, ECMP, VRF segmentation, across at least two major vendor platforms.
  • Deep RDMA expertise: lossless RoCE v2 tuning (PFC/ECN/DCQCN) and/or InfiniBand fabric management (subnet managers, adaptive routing, SHARP).
  • Hands-on experience with 100–800G optics, high-radix switching, and Clos/rail-optimized topologies.
  • Production network security experience: enterprise firewalls, site-to-site and remote-access VPNs, segmentation design.
  • Network automation proficiency: Python plus Ansible (or Nornir/NAPALM), with Git-based configuration workflows.
  • Working knowledge of the host-side RDMA stack (MOFED/DOCA, NIC tuning, GPUDirect) and how NCCL/RCCL collectives map onto physical fabric.
  • Cloud networking experience: VPC design and dedicated interconnects (Direct Connect/ExpressRoute class).
  • Dark fiber / DWDM procurement and operations experience.
  • BlueField DPU or SmartNIC deployments.
  • Experience supporting distributed training at 1,000+ GPU scale, or multi-cluster/cross-site training.
  • Kubernetes networking for GPU serving (CNI, SR-IOV, Multus).

Benefits

401(k)Vision Insurance