Principal Applied Scientist, Real-time Conversational AI , Agi

Amazon Amazon · Big Tech · Sunnyvale, CA · Applied Science

This role focuses on researching and developing real-time multimodal conversational AI, specifically advancing foundation models for speech and audio, and building post-training systems (reward modeling, RL) for natural conversational behavior. The scientist will contribute across the model lifecycle from pre-training to deployment, influencing technical direction and working with inference engineers for real-time production.

What you'd actually do

  1. Build and train large-scale multimodal foundation models for real-time speech and audio generation, from architecture design through production-scale training
  2. Design and build reward models and reward functions for speech systems — capturing naturalness, fluency, conversational quality, and real-time responsiveness
  3. Develop and apply reinforcement learning methods to shape conversational behavior — teaching models natural timing, responsiveness, and fluid interaction
  4. Design model architectures informed by hardware constraints and inference requirements, working with inference engineers to ensure models are servable from inception
  5. Design evaluation frameworks that capture the quality dimensions unique to real-time conversation

Skills

Required

  • PhD in Electrical Engineering, Computer Science, Mathematics, or a related technical field
  • 10+ years of industrial or academic work building speech recognition and natural language processing systems
  • Demonstrated experience training large-scale foundation models
  • Experience with multimodal model architectures that jointly process or generate across speech, text, and audio modalities
  • Strong understanding of transformer architectures and their application to speech/audio domains
  • Publication record at top-tier venues (NeurIPS, ICML, ICLR, Interspeech, ICASSP, ACL, or equivalent)
  • Demonstrated Principal/Staff+ scientific leadership

Nice to have

  • Hands-on experience building real-time AI systems — speech, audio, or video
  • Track record with post-training methods: reinforcement learning, reward modeling, RLHF/RLAIF, or alignment techniques applied to generative models
  • Experience with real-time interactive systems — models that handle concurrent input and output
  • Experience building speech-to-speech or audio-to-audio generative models
  • Hands-on design of reward models or reward functions specifically for speech/audio quality, naturalness, or conversational behavior
  • Experience with distributed training at scale
  • Familiarity with hardware-informed model design
  • Background in speech recognition, speech synthesis (TTS), or speech enhancement
  • Experience shipping research to production at scale
  • Contributions to open-source speech/audio ML systems or widely used research codebases
  • History of mentoring and growing research teams

What the JD emphasized

  • real-time
  • multimodal
  • speech
  • audio
  • foundation models
  • post-training
  • reward modeling
  • reinforcement learning
  • production deployment
  • scaling exercises
  • real-time latency budgets
  • shipping research to production at scale

Other signals

  • foundation models
  • real-time multimodal conversational AI
  • speech and audio generation
  • post-training systems
  • reward modeling
  • reinforcement learning
  • human-like conversational behavior
  • production deployment