AI Framework Engineer

AMD AMD · Semiconductors · Shanghai, China · Engineering

This role focuses on optimizing deep learning frameworks (TensorFlow, PyTorch) for AMD GPUs, enhancing GPU kernels, and improving training/inference performance on multi-GPU and multi-node systems. It involves working with internal GPU library teams and open-source maintainers, leveraging cutting-edge compiler technologies and advanced engineering principles.

What you'd actually do

  1. Optimize Deep Learning Frameworks: Enhance and optimize frameworks like TensorFlow and PyTorch for AMD GPUs in open-source repositories.
  2. Develop GPU Kernels: Create and optimize GPU kernels to maximize performance for specific AI operations.
  3. Develop & Optimize Models: Design and optimize deep learning models specifically for AMD GPU performance.
  4. Collaborate with GPU Library Teams: Work closely with internal teams to analyze and improve training and inference performance on AMD GPUs.
  5. Collaborate with Open-Source Maintainers: Engage with framework maintainers to ensure code changes are aligned with requirements and integrated upstream.

Skills

Required

  • C++ development
  • Linux environments
  • Python
  • deep learning frameworks (TensorFlow, PyTorch)
  • GPU kernel development
  • performance optimization
  • distributed computing environments
  • compiler technologies

Nice to have

  • HIP
  • CUDA
  • assembly (ASM)
  • AMD architectures (GCN, RDNA)
  • low-level programming
  • Compute Kernel (CK)
  • CUTLASS
  • Triton
  • LLVM
  • ROCm
  • heterogeneous compute clusters

What the JD emphasized

  • strong experience in C++ development within Linux environments
  • strong technical and analytical expertise
  • strong problem-solving skills
  • proactive approach
  • keen understanding of software engineering best practices
  • strong experience in designing and optimizing GPU kernels for deep learning on AMD GPUs using HIP, CUDA, and assembly (ASM)
  • Strong knowledge of AMD architectures (GCN, RDNA) and low-level programming to maximize performance for AI operations
  • Strong experience in integrating optimized GPU performance into machine learning frameworks (e.g., TensorFlow, PyTorch) to accelerate model training and inference, with a focus on scaling and throughput.
  • Expert skills in Python and C++
  • Strong experience in running large-scale workloads on heterogeneous compute clusters, optimizing for efficiency and scalability.
  • Sound understanding of compiler theory and tools like LLVM and ROCm for kernel and system performance optimization.

Other signals

  • optimizing deep learning frameworks
  • enhancing GPU kernels
  • training/inference performance
  • multi-GPU and multi-node systems
  • compiler technologies