Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

ML EngineerFull-timeLeadRemoteRemote

You will join a dedicated platform team to build a self-serve compute platform for GPU cluster infrastructure. You are required to have deep experience with Kubernetes and systems-level programming in Go, Rust, or C++. You will manage large-scale GPU fleets across multiple cloud providers, ensuring the reliability of both training and inference workloads. Expertise in distributed systems and high-speed networking like InfiniBand or RoCE is essential for this role.

Be the first to hear about new roles at Perplexity AI

Follow this company and we'll notify you as soon as new roles are posted.

Start free
Open