GPU Kernel Performance Engineer

AMD AMD · Semiconductors · Shanghai, China · Engineering

This role focuses on developing and optimizing ML kernels for GPU performance and efficiency, collaborating with cross-functional teams, and assisting in the design and implementation of ML algorithms and frameworks. It requires strong foundations in ML concepts, C++/Python, GPU programming (CUDA/ROCm), and ML frameworks like TensorFlow/PyTorch.

What you'd actually do

  1. Develop and optimize ML kernels with focus on performance and scalability.
  2. Collaborate with cross-functional teams to understand requirements and deliver high-quality software components.
  3. Assist in design and implementation of new ML algorithms and frameworks.
  4. Participate in code reviews, testing, and debugging processes to ensure the delivery of reliable software.
  5. Keep up-to-date with latest advancements in machine learning and hardware optimization techniques.

Skills

Required

  • C++
  • Python
  • GPU programming
  • parallel computing
  • TensorFlow
  • PyTorch

Nice to have

  • machine learning concepts
  • machine learning algorithms
  • CUDA
  • ROCm
  • problem-solving skills
  • attention to detail
  • team-oriented environment
  • communication skills

What the JD emphasized

  • ML kernels
  • performance and scalability
  • ML algorithms and frameworks
  • GPU programming
  • ML frameworks

Other signals

  • ML kernels
  • GPU programming
  • performance optimization