Research Scientist, AI Networking (phd)

Meta Meta · Big Tech · Menlo Park, CA

Research Scientist focused on optimizing the software stack for distributed ML training, specifically NCCL and PyTorch, to improve the reliability and performance of large-scale GenAI/LLM training and inference on Meta's GPU infrastructure.

What you'd actually do

  1. Enabling reliable and highly scalable distributed ML training on Meta's large-scale GPU training infra with a focus on GenAI/LLM scaling

Skills

Required

  • PhD degree in Computer Science, Computer Engineering, or relevant technical field
  • Specialized experience in High speed networking (RDMA), Distributed ML Training, GPU architecture, ML systems, AI infrastructure, high performance computing, performance optimizations, or Machine Learning frameworks (e.g. PyTorch)
  • Knowledge of ML, deep learning and LLM
  • Demonstrated software engineer experience
  • Experience in HPC and parallel computing
  • Knowledge of GPU architectures and CUDA programming
  • Experience with NCCL/RCCL/OneCCL and distributed GPU reliability/performance improvement on RoCE/Infiniband
  • Experience with both data parallel and model parallel training
  • Experience working and communicating cross-functionally in a team environment
  • Experience in AI framework and trainer development on accelerating large-scale distributed deep learning models
  • Experience working with DL frameworks like PyTorch, Caffe2 or TensorFlow

Nice to have

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • work authorization in country of employment

What the JD emphasized

  • PhD degree
  • publication track record
  • Experience in HPC and parallel computing
  • Experience with NCCL/RCCL/OneCCL and distributed GPU reliability/performance improvement on RoCE/Infiniband
  • Experience with both data parallel and model parallel training, such as Distributed Data Parallel, Fully Sharded Data Parallel (FSDP), Tensor Parallel, and Pipeline Parallel
  • Experience in AI framework and trainer development on accelerating large-scale distributed deep learning models

Other signals

  • distributed ML training
  • GPU communication stack
  • GenAI/LLM scaling reliability and performance