AI Training Optimization Engineer

AMD AMD · Semiconductors · Shanghai, China · Engineering

AMD is seeking an AI Training Optimization Engineer to help customers train AI models efficiently on AMD GPUs. The role involves identifying and filling gaps in the training ecosystem, optimizing critical kernels, and leveraging frontier techniques for training performance on large-scale systems. The engineer will also contribute to the development of kernel agents to accelerate kernel iteration and achieve extreme GPU performance.

What you'd actually do

  1. Ensure smooth training on AMD GPUs by identifying bottlenecks and delivering kernel-level performance improvements.
  2. Design and optimize kernels using HIP, CUDA, and Triton across real training workloads.
  3. Improve agent-based tooling to speed up kernel development and help achieve peak performance.
  4. Fill functional gaps, improve framework integration, and enhance ROCm-based training performance.
  5. Prototype next-generation kernels (e.g., sparse attention, linear attention ops).

Skills

Required

  • HIP
  • CUDA
  • Triton
  • GPU performance tuning
  • Transformer models
  • attention mechanisms
  • training algorithms
  • profiling and optimizing kernels
  • distributed training

Nice to have

  • PyTorch internals
  • Megatron-LM
  • DeepSpeed
  • kernel agents
  • runtime schedulers
  • performance-automation tools
  • kernel libraries (CUTLASS, CK)
  • Triton
  • ML compiler ecosystems

What the JD emphasized

  • kernel-level performance improvements
  • optimize kernels
  • kernel agents
  • training ecosystem
  • frontier kernel techniques

Other signals

  • optimizing critical kernels
  • frontier techniques to push the limits of training performance
  • kernel agents
  • AMD GPUs
  • large-scale systems