AI Software Engineer — High-performance GPU Communication

AMD AMD · Semiconductors · Shanghai, China · Engineering

AI Software Engineer focused on optimizing deep learning frameworks, GPU kernels, and training/inference performance for AMD GPUs in distributed environments. The role involves end-to-end optimization of inference and RL solutions, collaboration with internal library teams and open-source maintainers, and leveraging advanced compiler technologies.

What you'd actually do

  1. Build and optimize end to end distributed inference (e.g, P/D disaggregation and Large-EP) and RL solutions on mainstream frameworks like vLLM and SGlang.
  2. Work closely with internal teams to analyze and improve training and inference performance on AMD GPUs.
  3. Engage with framework maintainers to ensure code changes are aligned with requirements and integrated upstream.
  4. Optimize deep learning performance on both scale-up (multi-GPU) and scale-out (multi-node) systems.
  5. Leverage advanced compiler technologies to improve deep learning performance.

Skills

Required

  • C++ development
  • Linux environments
  • Python
  • software engineering best practices
  • deep learning frameworks
  • distributed computing environments
  • compiler technologies

Nice to have

  • GPU Kernel Development & Optimization
  • HIP
  • CUDA
  • assembly (ASM)
  • AMD architectures (GCN, RDNA)
  • low-level programming
  • Compute Kernel (CK)
  • CUTLASS
  • Triton
  • vLLM
  • SGlang
  • TensorFlow
  • PyTorch
  • LLM
  • multimodal
  • Text to Video
  • Image to Video
  • debugging
  • performance tuning
  • test design
  • High-Performance Computing
  • heterogeneous computing clusters
  • compiler theory
  • LLVM
  • ROCm

What the JD emphasized

  • End to end optimization
  • distributed inference
  • RL solutions
  • GPU Kernel Development & Optimization
  • Deep Learning Integration
  • End to end solution optimization
  • High-Performance Computing

Other signals

  • optimizing deep learning frameworks
  • enhancing GPU kernels
  • training/inference performance
  • distributed inference
  • RL solutions