Staff Software Development Engineer: Gpu, Computer Vision, Ai/ml Ops

AMD AMD · Semiconductors · Santa Clara, CA · Engineering

Staff Software Development Engineer at AMD focused on optimizing AI/ML performance on GPUs. This role involves architecting and driving the AI software stack, accelerating foundational models and AI agents, and innovating across hardware and software co-design. The engineer will bridge low-level GPU kernel engineering with AI post-training techniques and LLM inference optimization.

What you'd actually do

  1. Architect and Drive the AI Software Stack: You will establish best practices and optimize performance from the lowest-level GPU kernels to large-scale distributed systems, shaping the foundational software for AMD hardware. By leveraging cutting-edge Large Language Models (LLMs) and agent-based technologies, you will accelerate the development and performance enhancement of the AMD ROCm ecosystem, ensuring it remains at the forefront of AI innovation.
  2. Accelerate Foundational Models: Your work will directly accelerate cutting-edge applications like foundation models (LLMs) and autonomous AI agents, ensuring AMD is the platform of choice for the most demanding workloads.
  3. Innovate Across Hardware and Software: You will contribute to the entire co-design lifecycle, from influencing future GPU architectures to developing groundbreaking software for new accelerators and collaborating with the broader AI community.
  4. As a senior engineer, you will also be expected to mentor others and effectively communicate your ideas to shape the future of AI at AMD.

Skills

Required

  • high-performance C++ software engineering
  • low-level GPU programming
  • Large Language Models (LLMs)
  • AI systems
  • kernel engineering
  • AI post-training (RL)
  • GPU architectures (HIP/CUDA)
  • memory hierarchies
  • kernel optimization
  • large-scale C++/HIP/CUDA projects
  • ROCm ecosystem (e.g., rpp, MIVisionX, rocAL, rocdecode, rocjpeg)
  • CUDA libraries (e.g., CV-CUDA, cuDNN, NCCL)
  • C++/HIP/CUDA core of ML frameworks like PyTorch, TensorFlow, or JAX
  • transformer architectures
  • attention mechanisms
  • model lifecycle
  • Supervised Fine-Tuning (SFT)
  • Reinforcement Learning (e.g., RLHF, GRPO)
  • Mixture-of-Experts (MoE) architectures
  • inference optimizations (e.g., quantization, speculative decoding)
  • Agentic AI systems
  • code generation
  • self-improving LLMs
  • GPU programming (HIP/CUDA)
  • optimizing deep learning kernels and operators
  • GPU architecture
  • memory hierarchy
  • modern C++
  • object-oriented design
  • GPU profiling and performance analysis tools
  • distributed, multi-GPU systems
  • transformer architectures
  • attention mechanisms
  • Generative AI
  • Agentic AI
  • post-training pipelines of Large Language Models (LLMs)

Nice to have

  • Computer vision expertise
  • Experience or deep expertise with the AMD ROCm/HIP ecosystem.

What the JD emphasized

  • bridge kernel engineering with AI post-training (RL) experience
  • deep proficiency in high-performance C++ software engineering and low-level GPU programming
  • robust understanding of Large Language Models (LLMs) and AI systems
  • mastery in designing complex, scalable systems using modern C++
  • fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and kernel optimization
  • significant hands-on experience in large-scale C++/HIP/CUDA projects
  • deep understanding of LLMs, including but not limited to transformer architectures, attention mechanisms, and the full model lifecycle
  • hands-on experience in advanced model alignment and post-training techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning (e.g., RLHF, GRPO)
  • familiarity with cutting-edge trends such as Mixture-of-Experts (MoE) architectures, inference optimizations (e.g., quantization, speculative decoding), and modern application patterns like Agentic AI systems
  • Experience and interest in code generation and/or self-improving LLMs is a plus
  • Extensive hands-on experience in GPU programming (HIP/CUDA) and optimizing deep learning kernels and operators.
  • A fundamental understanding of GPU architecture and memory hierarchy, used to diagnose and resolve complex performance bottlenecks.
  • Deep experience using GPU profiling and performance analysis tools (e.g., AMD ROCm Profiler, NVIDIA Nsight) to diagnose and resolve complex bottlenecks in distributed, multi-GPU systems.
  • Deep knowledge of transformer architectures, attention mechanisms, and modern AI systems (Generative AI, Agentic AI).
  • Hands-on experience optimizing the post-training and inference pipelines of Large Language Models (LLMs).

Other signals

  • accelerate AI efficiency on GPUs
  • optimize performance from lowest-level GPU kernels to large-scale distributed systems
  • accelerate foundation models (LLMs) and autonomous AI agents
  • co-design lifecycle, from influencing future GPU architectures to developing groundbreaking software for new accelerators
  • bridge kernel engineering with AI post-training (RL) experience
  • optimize the post-training and inference pipelines of Large Language Models (LLMs)