Software Development Engineer I – Ai/ml Network Infrastructure, Annapurna Labs

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Software Development Engineer I role focused on the network infrastructure for EC2 distributed AI/ML systems, enabling large AI model training across GPU clusters. The role involves writing high-performance C/C++ code for communication libraries, building monitoring and automation infrastructure, and designing regression detection mechanisms.

What you'd actually do

  1. Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
  2. Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  3. Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  4. Design mechanisms to detect functional and performance regressions before they reach production
  5. Work across many instance types, software stacks, and Linux environments

Skills

Required

  • C/C++
  • Operating Systems (Linux internals, kernel concepts, memory management)
  • Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence)
  • Distributed Systems (consensus, message passing, fault tolerance, scalability)
  • Linux development environments and toolchains

Nice to have

  • ML communications
  • HPC networking
  • RDMA/high-speed interconnects
  • network programming (sockets, MPI, collective communication patterns)
  • performance profiling and optimization
  • GPU programming (CUDA)
  • hardware-software co-design
  • open-source projects in systems, networking, or HPC

What the JD emphasized

  • network stack for EC2 distributed AI/ML systems
  • enables the world's largest AI models to train
  • powering the largest AI workloads in the cloud

Other signals

  • enables the world's largest AI models to train
  • powering the largest AI workloads in the cloud
  • network stack for EC2 distributed AI/ML systems