Software Development Engineer, Ai/ml, Aws Neuron, Model Inference

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Software Development Engineer focused on optimizing AI/ML model inference performance on AWS custom hardware accelerators (Inferentia, Trainium) using the Neuron SDK. This role involves deep system-level and framework-level optimization, kernel development, and collaboration across hardware, software, and ML teams to ensure peak performance for LLMs and other GenAI workloads.

What you'd actually do

  1. Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  2. Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  3. Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
  4. Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models
  5. Analyze and optimize system-level performance across multiple generations of Neuron hardware

Skills

Required

  • Python
  • System level programming
  • ML knowledge
  • low-level optimization
  • system architecture
  • ML model acceleration
  • performance profiling
  • hardware-specific optimizations
  • distributed computing
  • PyTorch
  • JAX

Nice to have

  • compiler
  • runtime
  • frameworks
  • kernels
  • LLM optimization
  • latency optimization
  • throughput optimization

What the JD emphasized

  • critical to this role
  • must have
  • hard requirement

Other signals

  • AWS Neuron SDK
  • accelerate deep learning and GenAI workloads
  • Inferentia and Trainium ML accelerators
  • ML compiler, runtime, and application framework
  • PyTorch and JAX
  • ML inference and training performance
  • Inference Enablement and Acceleration team
  • novel architecture
  • maximizing their performance
  • PyTorch till the hardware-software boundary
  • systematic infrastructure
  • innovate new methods
  • high-performance kernels for ML functions
  • AI acceleration
  • frameworks and kernels
  • compiler to runtime and collectives
  • future architecture designs
  • intersection of machine learning, high-performance computing, and distributed architectures
  • AI acceleration technology
  • development, enablement and performance tuning
  • LLM model families
  • large language models
  • distributed inference solutions
  • optimizing inference performance for both latency and throughput
  • system level optimizations
  • Pytorch or JAX
  • distributed inference support for Pytorch
  • tune these models to ensure highest performance
  • efficiency of them running on the customer AWS Trainium and Inferentia silicon and servers
  • Python, System level programming and ML knowledge
  • low-level optimization, system architecture, and ML model acceleration
  • Design, develop, and optimize machine learning models and frameworks
  • custom ML hardware accelerators
  • distributed computing based architecture design
  • performance profiling
  • hardware-specific optimizations
  • production deployment
  • infrastructure to systematically analyze and onboard multiple models
  • high-performance kernels and features for ML operations
  • Neuron architecture and programming models
  • system-level performance
  • profiling tools
  • fusion, sharding, tiling, and scheduling
  • comprehensive testing
  • unit and end-to-end model testing
  • continuous deployment and releases through pipelines
  • enable and optimize their ML models
  • innovative optimization techniques
  • cross-functional team
  • applied scientists, system engineers, and product managers
  • state-of-the-art inference capabilities
  • Generative AI applications
  • debugging performance issues
  • optimizing memory usage
  • shaping the future of Neuron's inference stack
  • Amazon and the Open Source Community
  • drive efficiencies in software architecture
  • create metrics
  • implement automation
  • resolve the root cause of software defects
  • high-impact solutions