Sr. Software Development Engineer, Hpc/ml Networking Engineer, Annapurna Labs

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

This role focuses on distributed AI/ML systems, specifically collective operations that enable AI to scale across multiple accelerators and servers. The work involves C/C++ and low-level systems programming, with a focus on Linux, kernels, and performant code. Experience with embedded systems, high-speed networking, or HPC interconnects is valued. The role is part of Annapurna Labs, contributing to EC2 infrastructure for large-scale AI models and customers.

What you'd actually do

  1. work on distributed AI/ML systems
  2. working on collective operations - the fundamental operations that enable AI to scale across multiple accelerators & servers
  3. solid knowledge of Linux, kernels, and performant code is important
  4. Experience with embedded systems is valued, and experience with high-speed networking or HPC interconnects is valued highly
  5. work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale

Skills

Required

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience as a mentor, tech lead or leading an engineering team
  • C/C++ Coding Experience

Nice to have

  • Bachelor's degree in computer science or equivalent
  • ML Communications (NCCL, NIXL, NVSHMEM)
  • embedded systems
  • high-speed networking
  • HPC interconnects

What the JD emphasized

  • Must have C/C++ Coding Experience

Other signals

  • distributed AI/ML systems
  • collective operations
  • scale across multiple accelerators & servers
  • largest clusters, with the largest customers, for the largest AI models