Sr. Software Development Engineer - Collectives and Network

AMD AMD · Semiconductors · Austin, TX · Engineering

This role focuses on optimizing AI pre-training and distributed inference performance on AMD GPUs. The engineer will work on strategy, architecture, optimization, and tooling across the software stack, including hardware-software co-design, performance tuning, profiling, and developing estimation tools. The role requires deep knowledge of network, NIC, and GPU hardware, software optimization, AI frameworks, and distributed systems for large-scale models.

What you'd actually do

  1. Performance tuning, profiling and analysis of large-scale models for LLM, diffusion, multimodal, RecSys and generative AI, single node and distributed. In addition to exploring various tradeoffs and design decisions.
  2. Develop and improve framework, tools and infrastructure for performance estimation, modeling and reporting.
  3. Participate in hardware-software co-design for future hardware optimizations – especially on scale-up networks, NIC and scale-out networks.
  4. Provide guidelines to customers on efficient network load-balancing, workload scheduling and model sharding strategies.
  5. Help with strategy and roadmap for AMD Collectives and Network optimizations.

Skills

Required

  • Network, NIC and GPU hardware architecture
  • software optimization
  • performance modeling
  • AI frameworks
  • inference and training optimization
  • mapping model architecture to low level software
  • distributed inference
  • PyTorch
  • JAX
  • vLLM
  • SGLang

Nice to have

  • technical leadership skills
  • work collaboratively with cross-functional teams
  • Mentor, coach, and inspire a diverse and talented team of researchers and engineers
  • Excellent written, verbal, and presentation skills
  • coordinate internally and externally

What the JD emphasized

  • latest state-of-the-art AI models
  • distributed inference and deployment at scale is crucial
  • performance optimization
  • scale-up networks, NIC and scale-out networks

Other signals

  • performance optimization
  • distributed inference
  • AI pre-training
  • AMD GPU
  • ROCm