Sr. Sde, Edge AI ML Platform, Edge AI and Science

Amazon Amazon · Big Tech · CA, BC +1 · Software Development

Senior Software Development Engineer to lead the architecture and delivery of core ML platform capabilities for training, optimizing, evaluating, and deploying generative AI models on devices and in the cloud. The role involves solving problems across distributed training, model onboarding, compression, evaluation, GPU performance, artifact management, CI/CD, observability, and operational reliability for large-scale models.

What you'd actually do

  1. Lead the design and delivery of distributed ML platform services and libraries across model ingestion, optimization, training, evaluation, packaging, and deployment.
  2. Define stable APIs and architecture boundaries that allow scientists to add algorithms without coupling research code to training, infrastructure, or deployment implementations.
  3. Design distributed training capabilities across data, tensor, pipeline, and model parallelism for large language and multimodal models.
  4. Scale workflows on multi-node GPU clusters while improving training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration time.
  5. Develop infrastructure that connects distributed training with distillation, quantization, pruning, and other model optimization techniques.

Skills

Required

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Bachelor's degree in Computer Science, Engineering, or a related technical field
  • Experience designing or building distributed systems or high-performance computing systems.

Nice to have

  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience with CUDA kernels or ML/low-level kernels, or experience in debugging, profiling, and implementing software engineering best practices in large-scale systems
  • Experience programming with at least one modern language such as Jav

What the JD emphasized

  • core ML platform capabilities
  • distributed training
  • model onboarding
  • compression
  • evaluation
  • GPU performance
  • artifact management
  • CI/CD
  • observability
  • operational reliability
  • hundreds of billions of parameters
  • technical leadership
  • resolve ambiguous requirements
  • lead projects that span multiple engineers and teams
  • raise the engineering bar
  • model ingestion
  • optimization
  • training
  • evaluation
  • packaging
  • deployment
  • stable APIs
  • architecture boundaries
  • research code
  • training
  • infrastructure
  • deployment implementations
  • distributed training capabilities
  • data parallelism
  • tensor parallelism
  • pipeline parallelism
  • model parallelism
  • large language
  • multimodal models
  • multi-node GPU clusters
  • training throughput
  • GPU utilization
  • memory efficiency
  • communication performance
  • failure recovery
  • developer iteration time
  • distributed training
  • distillation
  • quantization
  • pruning
  • model optimization techniques
  • evaluation
  • artifact workflows
  • model quality
  • system performance
  • validated models
  • deployment on target hardware
  • automated validation
  • CI/CD
  • regression testing
  • observability
  • release mechanisms
  • GPU-intensive ML workloads
  • Profile and optimize end-to-end system performance
  • applied scientists
  • GPU kernel engineers
  • bottlenecks
  • durable platform improvements
  • operational mechanisms
  • metrics
  • alarms
  • runbooks
  • on-call practices
  • root-cause correction
  • production platform services
  • model
  • compiler
  • runtime
  • hardware
  • security
  • infrastructure teams
  • technical dependencies
  • multi-team programs
  • technical designs
  • evaluate trade-offs
  • build consensus
  • customer need
  • technology strategy
  • Mentor engineers
  • code and design review practices
  • recruit and develop a strong engineering team
  • architecture
  • implementation
  • model onboarding interfaces
  • failures in distributed training runs
  • profiling GPU workloads
  • cross-team reviews
  • end-to-end deployment paths
  • platform abstractions
  • release and regression mechanisms
  • model teams
  • performance
  • reliability
  • developer productivity data
  • prioritize platform investments
  • incremental deliveries
  • long-term architecture
  • team resolves recurring problems
  • root
  • reusable model training
  • optimization
  • deployment capabilities
  • Amazon product teams
  • applied scientists
  • Edge AI
  • adapt rapidly changing model architectures
  • constrained hardware
  • production workloads
  • rebuilding the toolchain
  • every model
  • platform foundations
  • model development
  • deployment
  • end-to-end scope
  • training
  • compression
  • evaluation
  • deployment as one system
  • clear interfaces
  • measurable performance
  • automated quality gates
  • direct collaboration
  • science and engineering
  • non-internship professional software development experience
  • programming with at least one software programming language experience
  • leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • mentor
  • tech lead
  • leading an engineering team
  • Computer Science, Engineering, or a related technical field
  • designing or building distributed systems or high-performance computing systems
  • full software development life cycle
  • coding standards
  • code reviews
  • source control management
  • build processes
  • testing
  • operations experience
  • CUDA kernels
  • ML/low-level kernels
  • debugging
  • profiling
  • implementing software engineering best practices
  • large-scale systems
  • programming with at least one modern language such as Jav

Other signals

  • ML platform
  • distributed training
  • model optimization
  • deployment on devices