Research Engineer, Generative Media, Deepmind

Google Google · Big Tech · Tel Aviv, Israel

Research Engineer role focused on building frontier generative media experiences, including real-time audio-visual conversations and video stream manipulation. The role involves designing, implementing, training, and evaluating novel deep learning models for media synthesis and manipulation, with a focus on live dialog, live video editing, and crosscut applications. Collaboration with research scientists and engineers is key, along with contributing to datasets, evaluation methodologies, and infrastructure. The role also involves analyzing results for publications and product deployment.

What you'd actually do

  1. Design, implement, train, and evaluate novel deep learning models for real-time generative media, including video-audio synthesis, manipulation, and understanding.
  2. Develop and optimize algorithms and systems for AI experiences on Omni platforms, focusing on live dialog, live video editing, and crosscut.
  3. Collaborate with research scientists and engineers to prototype, experiment, and scale generative AI techniques using software engineering best practices.
  4. Contribute to the design of datasets, evaluation methodologies, and infrastructure required for training and testing generative models.
  5. Analyze results to iterate on architectures, contribute to research publications, and help deploy successful research into GDM and Google products.

Skills

Required

  • Designing, training, and evaluating deep generative models (e.g., Diffusion Models, Transformers, GANs) for media synthesis and manipulation (image, video, audio)
  • Experimentation, dataset curation, eval systems, and building ML pipelines
  • Machine learning, computer vision, and deep learning
  • Python and deep learning frameworks such as JAX, TensorFlow, or PyTorch

Nice to have

  • Multimodal learning, integrating video, audio, and text
  • Video generation and editing models
  • Knowledge of techniques for model optimization and efficient/low-latency inference suitable for real-time applications
  • Familiarity with real-time systems or media streaming technologies

What the JD emphasized

  • real-time generative media
  • video-audio synthesis
  • manipulation and enhancement of video streams
  • foundational generative models
  • interactive experiences
  • real-time audio-visual conversations
  • instantaneous, on-the-fly manipulation and enhancement of video streams
  • real-time generative media
  • video-audio synthesis
  • manipulation and understanding
  • live dialog
  • live video editing
  • crosscut
  • generative AI techniques
  • datasets
  • evaluation methodologies
  • infrastructure
  • training and testing generative models
  • research publications
  • deploy successful research
  • deep generative models
  • media synthesis and manipulation
  • experimentation
  • dataset curation
  • eval systems
  • ML pipelines
  • machine learning
  • computer vision
  • deep learning
  • multimodal learning
  • video generation and editing models
  • model optimization
  • efficient/low-latency inference
  • real-time applications
  • real-time systems
  • media streaming technologies

Other signals

  • real-time generative media
  • video-audio synthesis
  • manipulation and enhancement of video streams
  • foundational generative models
  • interactive experiences