Senior Product Manger - Tech, Infrastructure Reliability

Amazon Amazon · Big Tech · Austin, TX · Project/Program/Product Management--Technical

Product Manager role focused on building an AI-powered platform for Amazon's fulfillment network, using LLMs and multi-agent systems for incident detection, diagnosis, and resolution. The role involves roadmap ownership, writing code, building proof-of-concepts, and partnering with data scientists and engineers on model architecture and agent orchestration.

What you'd actually do

  1. You will own and drive the multi-year product roadmap for the Infrastructure Reliability AI-Ops platform across three strategic programs: zero-touch incident resolution, associate-directed work tooling, and predictive failure prevention.
  2. Beyond traditional product management, you will write code and deliver working proof-of-concepts — prototyping multi-agent reasoning pipelines, testing anomaly detection approaches, and stress-testing LLM prompt chains against real incident data — to validate technical hypotheses before committing engineering resources.
  3. You will apply deep machine learning knowledge to shape how the platform detects, consolidates, and reasons about failures, engaging meaningfully with data scientists on model architecture selections, feature engineering tradeoffs, and evaluation frameworks.
  4. You will define how the platform uses chain-of-thought prompting, retrieval-augmented generation, and confidence calibration to build progressive confidence about incident severity and origin, and you will define the multi-agent architecture coordinating detection, investigation, diagnosis, and remediation, including agent roles, handoff conditions, and safety boundaries.
  5. You will translate operational needs into a prioritized backlog, representing Incident Managers, domain engineers, and Operations Control Center stakeholders during executive-level planning, while defining and tracking the business case and driving cross-functional alignment across Fulfillment Technologies, Robotics, and Network Engineering.

Skills

Required

  • technical product management
  • creating written docs for development of new products
  • enterprise security product experience
  • product management in the cloud computing technology space
  • full product life cycle experience
  • technical (software development, network development, IT, other related) experience
  • creating written docs for development of new products

Nice to have

  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience working within teams delivering software products

What the JD emphasized

  • multi-year roadmap
  • write code
  • build proof-of-concepts
  • multi-agent reasoning pipelines
  • anomaly detection approaches
  • LLM prompt chains
  • model architecture
  • multi-agent orchestration
  • zero-touch incident resolution
  • predictive failure prevention
  • multi-agent architecture

Other signals

  • LLM
  • multi-agent systems
  • anomaly detection
  • predictive failure prevention
  • zero-touch incident resolution