Software Development Engineer Ai/ml, Inference Model Enablement, Aws Neuron

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Software Development Engineer role focused on optimizing and enabling state-of-the-art LLMs for inference on AWS Neuron custom hardware accelerators. Responsibilities include performance optimization, system reliability, driving architectural decisions for the serving stack, and collaborating with cross-functional teams. Experience with distributed inference libraries and ML/LLM fundamentals is preferred.

What you'd actually do

  1. Deliver high-performance models using distributed inference libraries
  2. Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem
  3. Mentor team members and provide technical leadership across multiple work streams
  4. Drive architectural decisions that impact the entire Neuron serving stack
  5. Collaborate with customers, product owners, and engineering teams to define technical strategy

Skills

Required

  • 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • 5+ years of non-internship professional software development experience
  • Experience as a mentor, tech lead or leading an engineering team

Nice to have

  • Master's degree in computer science or equivalent
  • Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution

What the JD emphasized

  • optimize the latest models to run really fast on the Trainium hardware
  • onboard and optimize state-of-the-art open-source and customer LLMs
  • advancing inference usability and quality
  • state-of-the-art inference capabilities

Other signals

  • optimizing LLMs for inference on custom hardware accelerators
  • improving model enablement speed and experience
  • advancing inference usability and quality