Software Development Engineer, Aws Marketplace & Partner Services

Amazon Amazon · Big Tech · Seattle, WA · Software Development

Software Engineer role focused on building and scaling the infrastructure and tooling for AI solutions within AWS Marketplace, including LLM training, model hosting, and AI agent systems. The role involves designing, developing, and operating scalable systems for AI agent performance monitoring, evaluation, and observability, as well as contributing to research initiatives.

What you'd actually do

  1. Design, develop, and maintain scalable infrastructure for large language model training pipelines, model hosting, and agent-based systems, employing best engineering practices and iterating from simple to complex solutions
  2. Build and operate monitoring, evaluation, and observability frameworks to measure AI agent performance, reliability, and effectiveness in production
  3. Develop reusable tooling, APIs, and automation that accelerate the team's ability to experiment, deploy, and iterate on ML/AI solutions
  4. Contribute to a fast-paced, experimental environment by rapidly prototyping and operationalizing infrastructure that leverages the most recent advances in ML/AI
  5. Collaborate with cross-functional teams—including Data Scientists, Applied Scientists, Data Engineers, Software Engineers, Product Managers, and Science and Engineering leaders—to translate scientific innovations into robust, production systems

Skills

Required

  • Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language
  • 2+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services
  • 2+ years of experience working with cloud services (e.g., AWS, Azure, GCP)

Nice to have

  • Experience with LLM training pipelines
  • Experience with custom LLM model hosting
  • Experience with AI agent monitoring and evaluation infrastructure
  • Experience with AI agents
  • Experience with AWS scale
  • Experience with production systems

What the JD emphasized

  • scalable infrastructure
  • large language model training pipelines
  • custom LLM model hosting
  • AI agent monitoring and evaluation infrastructure
  • AI agents
  • AWS scale
  • production systems

Other signals

  • LLM training pipelines
  • custom LLM model hosting
  • AI agent monitoring and evaluation infrastructure
  • AWS scale
  • production systems