Sr. Sdm, AI Inference, Neuron Sdk

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

This role leads a team of managers and engineers to optimize and deploy AI inference models on Amazon's custom ML accelerators (Trainium, Inferentia). Responsibilities include the full development lifecycle of model onboarding, low-level performance optimization, feature support, and efficient serving, reliability, and scalability. The role requires deep understanding of a vertically integrated system stack from hardware to serving.

What you'd actually do

  1. lead a strong team of managers and engineers to optimize models for best performance and robustness when deployed at scale on Trainium and Inferentia devices
  2. responsible for the full development life cycle of model onboarding, low-level performance optimization, feature support and efficient serving, including reliability and scalability
  3. work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers
  4. build massive-scale distributed training and inference solutions, developing the full stack of software, servers and chips together with teams across the Annapurna organization to run the largest machine learning workloads

Skills

Required

  • 10+ years of engineering experience
  • 5+ years of engineering team management experience
  • 10+ years of planning, designing, developing and delivering consumer software experience
  • Experience partnering with product or program management teams
  • Experience managing multiple concurrent programs, projects and development teams in an Agile environment

Nice to have

  • Experience partnering with product and program management teams
  • Experience designing and developing large scale, high-traffic applications

What the JD emphasized

  • optimize models for best performance and robustness
  • low-level performance optimization
  • efficient serving, including reliability and scalability
  • optimizing and serving AI models
  • vertically integrated system stack that consisting of hardware, frameworks, serving stack, model and workflows

Other signals

  • optimize and deploy inference models
  • optimize models for best performance and robustness
  • efficient serving, including reliability and scalability
  • optimizing and serving AI models
  • vertically integrated system stack that consisting of hardware, frameworks, serving stack, model and workflows