ML Software Engineer, Data Plane

Amazon Amazon · Big Tech · Cupertino, CA · Software Development

ML Software Engineer focused on the inference data plane for large models on custom hardware, optimizing kernels, memory management, data movement, and serving integration for production-level performance.

What you'd actually do

  1. Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  2. Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  3. Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  4. Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  5. Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up.

Skills

Required

  • 4+ years of full software development life cycle
  • Knowledge of computer architecture, operating systems, and parallel computing
  • Knowledge of Linux fundamentals
  • Strong proficiency in C/C++
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators

Nice to have

  • Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
  • Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT
  • Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware
  • Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming
  • Experience with hardware simulation environments and model validation workflows
  • Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow

What the JD emphasized

  • own the design and implementation
  • write and optimize low-level code
  • validate model architectures end-to-end
  • drive performance across the stack
  • Proven track record of owning and delivering complex software features end-to-end

Other signals

  • custom hardware
  • large distributed models
  • low-level code optimization
  • inference performance