Sr. Software Engineer - AI Triton Communication

AMD AMD · Semiconductors · San Jose, CA · Engineering

This role focuses on advancing the Triton language and compiler for AMD GPUs, specifically building native distributed execution and communication capabilities to enable efficient large-scale AI training and inference. The engineer will work across compiler, runtime, and hardware layers to optimize performance and scalability for AI workloads on AMD Instinct Accelerators.

What you'd actually do

  1. Design and develop native distributed communication and execution capabilities within the Triton AMDGPU backend, enabling scalable multi-GPU execution for large-scale AI workloads
  2. Design and implement Triton compiler and runtime mechanisms for native GPU-initiated communication, including collective operations, remote memory access, synchronization, and distributed execution primitives
  3. Drive performance optimization across compute and communication, including inter-GPU data movement, communication/computation overlap, memory hierarchy utilization, and GPU-driven scheduling efficiency
  4. Develop and optimize distributed Triton kernels and execution models to achieve high performance, scalability, and efficient hardware utilization for AI workloads
  5. Analyze, profile and debug complex cross-stack issues spanning Triton compiler, runtime, ROCm stack, and GPU hardware execution

Skills

Required

  • GPU architecture
  • compiler technologies
  • distributed GPU systems
  • GPU programming
  • performance optimization
  • system-level performance challenges
  • Triton compiler
  • runtime infrastructure
  • multi-GPU scale
  • cross-stack issues analysis
  • collaboration with GPU architecture, compiler, runtime, and performance teams

Nice to have

  • Triton compiler and runtime
  • modern GPU architectures
  • execution model
  • memory hierarchy
  • scheduling
  • occupancy
  • hardware performance characteristics
  • GPU runtime systems
  • communication stacks
  • multi-GPU interconnects
  • distributed GPU communication libraries
  • RCCL
  • NCCL
  • NVSHMEM
  • rocSHMEM
  • MPI
  • inter-GPU communication
  • synchronization
  • communication/computation overlap
  • MLIR
  • LLVM internals
  • profiling
  • debugging
  • ROCm
  • HIP
  • CUDA
  • GPU programming ecosystems
  • performance profiling and optimization tools
  • large-scale AI, machine learning or HPC workloads
  • open-source projects
  • collaborative, cross-functional engineering environments
  • problem-solving
  • communication
  • technical leadership

What the JD emphasized

  • first-class support for distributed execution and communication in Triton is strategically critical
  • delivering industry-leading distributed performance and scalability on AMD Instinct accelerators is a key priority
  • performance, scalability, and usability of Triton directly impact the competitiveness of AMD hardware in large-scale AI deployments
  • deep expertise in GPU architecture, compiler technologies, and distributed GPU systems
  • proven experience optimizing workloads at multi-GPU scale
  • comfortable working across the full execution stack — from compiler and runtime to hardware
  • experience working close to the GPU runtime, communication stack, or compiler backend
  • motivated to build native distributed execution and communication capabilities tightly integrated with the compiler and runtime to maximize scalability and hardware utilization
  • thrive on solving complex system-level performance challenges and delivering scalable, high-performance GPU infrastructure

Other signals

  • enabling efficient large-scale training and inference on AMD Instinct Accelerators
  • building native distributed execution and communication capabilities
  • performance optimization across compute and communication
  • contribute to open-source Triton and ROCm distributed ecosystem