Senior ML Software Engineer, Data Plane

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

Senior ML Software Engineer focused on designing and implementing the inference data plane for large models on custom hardware, optimizing model execution, memory management, data movement, and serving integration for production-level performance.

What you'd actually do

  1. Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  2. Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  3. Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  4. Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  5. Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.

Skills

Required

  • C/C++
  • Linux systems knowledge
  • Developing compute kernels for GPUs, DSPs, or custom accelerators
  • Owning and delivering complex software features end-to-end

Nice to have

  • Machine Learning and LLM fundamentals
  • Transformer architecture
  • Training/inference lifecycles
  • Optimization techniques
  • ML frameworks (JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT)
  • Speculative decoding
  • KV cache optimization
  • LLM serving optimizations
  • Distributed systems
  • Hardware simulation environments
  • Model validation workflows
  • AI-assisted development tools

What the JD emphasized

  • production-level performance
  • custom hardware
  • large language model inference
  • end-to-end
  • custom accelerator backends
  • model parallelism
  • hardware targets
  • hardware bringup
  • low-level code
  • custom ML accelerator architecture

Other signals

  • custom hardware
  • large distributed models
  • low-level code optimization
  • inference performance