Software Development Engineer, Annapurna Labs, Elastic Collectives

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Software Development Engineer role focused on distributed AI/ML systems, specifically collective operations for scaling AI across multiple accelerators and servers. The role involves low-level C/C++ development on Linux, with a focus on performance, kernels, and potentially embedded systems, high-speed networking, or HPC interconnects. It's positioned at the forefront of AI/ML, supporting large-scale AI models and customers within Annapurna Labs, an integral part of AWS EC2 infrastructure.

What you'd actually do

  1. work on collective operations - the fundamental operations that enable AI to scale across multiple accelerators & servers
  2. work on features for the largest clusters, with the largest customers, for the largest AI models
  3. building networking solutions that for Machine Learning (ML) and High-Performance Computing (HPC) workloads on AWS
  4. working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects
  5. mentoring new and junior engineers

Skills

Required

  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language
  • Experience with C/C++

Nice to have

  • Experience with embedded systems
  • experience with high-speed networking or HPC interconnects
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor's degree in computer science or equivalent

What the JD emphasized

  • solid knowledge of Linux, kernels, and performant code is important
  • Experience with embedded systems is valued
  • experience with high-speed networking or HPC interconnects is valued highly

Other signals

  • distributed AI/ML systems
  • collective operations
  • scale across multiple accelerators & servers
  • largest clusters
  • largest customers
  • largest AI models