Software Development Manager, LLM Inference Model Enablement, Neuron Sdk

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Software Development Manager to lead a team optimizing LLMs for inference on AWS Neuron custom hardware accelerators, focusing on performance, enablement speed, and inference usability.

What you'd actually do

  1. Management and execution against project plans and delivery commitments
  2. Manage the day-to-day activities of the engineering team
  3. Management of resources, staffing, mentoring, and maintaining a best-of-class engineering team
  4. Report on status of development, quality, operations, and model performance to management

Skills

Required

  • 3+ years of engineering team management experience
  • 7+ years of working directly within engineering teams experience
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Experience partnering with product or program management teams

Nice to have

  • Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
  • Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers
  • LLM model architectures
  • model performance optimizations
  • inference techniques
  • delivering high-performance models using distributed inference libraries
  • managing demanding, fast-changing priorities
  • technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives

What the JD emphasized

  • optimize LLMs to run really fast on the Trainium hardware
  • state-of-the-art open-source and customer LLMs
  • inference on Trainium accelerators
  • model enablement speed and experience
  • inference usability and quality
  • high-performance models using distributed inference libraries
  • PyTorch inference library, Neuron compiler, runtime, and collectives

Other signals

  • optimizing LLMs for inference on custom hardware accelerators
  • improving model enablement speed and experience
  • advancing inference usability and quality