Senior Software Development Engineer, Aws Mantle

Amazon Amazon · Big Tech · Seattle, WA · Software Development

Senior Software Development Engineer on the AWS Mantle team, responsible for the technical vision and end-to-end system design of a distributed inference engine that serves foundation models for Amazon Bedrock. Focuses on latency, reliability, scalability, and security at a massive scale.

What you'd actually do

  1. Set the long-term technical direction for a globally distributed, high-performance ML inference platform serving models from industry-leading AI providers
  2. Own end-to-end system design decisions that directly impact latency, reliability, and scalability for millions of customers worldwide
  3. Influence engineering strategy across Amazon Bedrock, partnering with senior leadership to align technical investments with business outcomes
  4. Raise the engineering bar through exemplary system design, mentorship, and contributions to the broader AWS engineering community
  5. Navigate complex trade-offs across performance, security, and cost while maintaining the highest standards for operational excellence

Skills

Required

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Experience driving cross-organizational technical strategy and delivering results in complex, ambiguous environments

Nice to have

  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Master's degree or equivalent in computer science, machine learning, engineering, or related fields, or PhD
  • Experience building large-scale machine learning and AI solutions at Internet scale
  • Familiarity with inference frameworks such as vLLM, TensorRT, or Triton Inference Server

What the JD emphasized

  • architect solutions that are reliable, scalable, and secure
  • millisecond-level latency
  • zero-trust security
  • scaling inference
  • global availability and performance SLAs
  • secure, enterprise-grade access
  • Zero Operator Access architecture

Other signals

  • distributed inference engine
  • serving foundation models
  • AWS Bedrock
  • millisecond-level latency
  • security