AI Framework Engineer

AMD AMD · Semiconductors · Shanghai, China · Engineering

This role focuses on optimizing deep learning frameworks and performance for AMD GPUs, specifically enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and multi-node systems. The engineer will work on optimizing distributed inference and RL solutions using frameworks like vLLM and SGlang, collaborating with internal GPU library teams and open-source maintainers, and leveraging compiler technologies.

What you'd actually do

  1. Build and optimize end to end distributed inference (e.g, P/D disaggregation and Large-EP) and RL solutions on mainstream frameworks like vLLM and SGlang.
  2. Work closely with internal teams to analyze and improve training and inference performance on AMD GPUs.
  3. Engage with framework maintainers to ensure code changes are aligned with requirements and integrated upstream.
  4. Optimize deep learning performance on both scale-up (multi-GPU) and scale-out (multi-node) systems.
  5. Leverage advanced compiler technologies to improve deep learning performance.

Skills

Required

  • C++ development
  • Linux environments
  • Python
  • software engineering best practices
  • debugging
  • performance tuning
  • test design

Nice to have

  • GPU Kernel Development & Optimization
  • HIP
  • CUDA
  • assembly (ASM)
  • AMD architectures (GCN, RDNA)
  • low-level programming
  • Compute Kernel (CK)
  • CUTLASS
  • Triton
  • Deep Learning Integration
  • vLLM
  • SGlang
  • TensorFlow
  • PyTorch
  • LLM and multimodal
  • distributed inference
  • RL
  • Text to Video
  • Image to Video
  • High-Performance Computing
  • heterogeneous computing clusters
  • Compiler Optimization
  • LLVM
  • ROCm

What the JD emphasized

  • optimizing deep learning frameworks
  • enhancing GPU kernels, deep learning models, and training/inference performance
  • multi-GPU and multi-node systems
  • optimizing distributed inference
  • RL solutions on mainstream frameworks like vLLM and SGlang
  • GPU Kernel Development & Optimization
  • Deep Learning Integration
  • End to end solution optimization
  • Compiler Optimization

Other signals

  • optimizing deep learning frameworks
  • enhancing GPU kernels, deep learning models, and training/inference performance
  • multi-GPU and multi-node systems
  • optimizing distributed inference
  • RL solutions on mainstream frameworks like vLLM and SGlang