Fellow Software Engineer — AI Performance & Reliability

AMD AMD · Semiconductors · San Jose, CA · Engineering

Fellow Software Engineer focused on AI Performance & Reliability, optimizing AI workloads (training and inference) for LLMs, diffusion models, and recommendation systems. The role involves profiling, identifying bottlenecks across the stack (models, frameworks, hardware), developing performance tooling, and collaborating with customers and internal teams to improve efficiency, latency, throughput, and reliability. Requires strong software engineering, systems performance, and ML framework experience.

What you'd actually do

  1. Profile and optimize AI model training and inference workloads.
  2. Improve model throughput, latency, memory efficiency, scalability, and reliability.
  3. Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  4. Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
  5. Develop performance tooling, benchmarks, automation, and observability systems.

Skills

Required

  • Strong software engineering skills and experience building production-quality systems.
  • Experience working with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication.
  • Proficiency in languages such as Python, C++, or similar systems-oriented programming languages.
  • Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack.
  • Clear written and verbal communication skills.
  • A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.

Nice to have

  • Experience optimizing large language models, diffusion models, or recommendation models.
  • Experience with GPU, accelerator, or distributed computing environments.
  • Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.
  • Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.
  • Experience operating AI systems in production environments.
  • Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role.
  • Experience designing benchmarks and conducting systematic performance analysis.

What the JD emphasized

  • improving the performance, efficiency, and reliability of AI workloads across both model training and inference
  • optimize workloads involving large language models, diffusion models, recommendation systems
  • develop performance tooling, benchmarks, automation, and observability systems
  • customer-focused mindset and willingness to work directly with customers

Other signals

  • improving the performance, efficiency, and reliability of AI workloads across both model training and inference
  • optimize workloads involving large language models, diffusion models, recommendation systems
  • develop performance tooling, benchmarks, automation, and observability systems
  • optimize AI model training and inference workloads
  • improve model throughput, latency, memory efficiency, scalability, and reliability