Research Engineer, Large-scale Training

Together AI Together AI · Data AI · San Francisco, CA · Research

Research Engineer focused on scaling and optimizing large-scale training infrastructure for foundation models, integrating new architectures, and productionizing novel training methods.

What you'd actually do

  1. Design, implement, and optimize core components of Together's large-scale training infrastructure.
  2. Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
  3. Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
  4. Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
  5. Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.

Skills

Required

  • Python
  • PyTorch
  • training or fine-tuning large neural networks
  • multi-GPU or multi-node environments
  • ML systems fundamentals
  • GPU architecture
  • mixed-precision training
  • distributed training paradigms

Nice to have

  • CUDA
  • Triton
  • NCCL
  • NVSHMEM
  • FSDP
  • DeepSpeed
  • Megatron-LM
  • custom distributed training systems
  • compute efficiency optimization
  • memory efficiency optimization
  • scalability optimization
  • large-scale GPU experiments
  • open-source ML projects
  • ML products or managed services

What the JD emphasized

  • Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
  • Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.

Other signals

  • large-scale training infrastructure
  • efficient model training
  • production fine-tuning workloads
  • new methods for efficient model training and evaluation