Member of Technical Staff, Agentic Systems - Games

Netflix Netflix · Big Tech · Los Gatos, CA +2 · Data & Insights

Netflix is seeking a technical leader for their Games division to shape the strategy, development, and delivery of GenAI tools and agentic systems. This role involves hands-on engineering to build multi-step reasoning pipelines, tool-use agents, and multi-agent orchestration, as well as defining product strategy, working across teams, and managing evaluation infrastructure. The position requires experience in AI product strategy, machine learning, AI engineering, and game development, with a focus on shipping production-grade agentic systems.

What you'd actually do

  1. Builder PM: Identify key opportunities where AI can add value to teams, defining the product strategy and architecture of new enablement systems.
  2. Work across diverse teams in game studios and platforms, to inform and advise on existing initiatives, as well as how new tools will fit within the broader ecosystem.
  3. Contribute to engineering production-grade agentic systems: multi-step reasoning pipelines, tool-use agents, multi-agent orchestration, and autonomous workflows. Design and build the code harnesses and scaffolding connecting frontier models (or open-weight alternatives) to tooling, game engines, and platform APIs.
  4. Data Driven Decision Making: Define AI evaluation as a first-class discipline, including defining overall product evaluations strategy, curating offline eval sets, automated scoring pipelines, and regression gates. Define and build online evaluation: A/B testing, production telemetry, user feedback loops, and anomaly detection for agent behavior drift.
  5. Collaborate with our engineering and platform teams to build robust solutions and scale core capabilities: model inference, data pipelines, responsible AI compliance, safety guardrails, and graceful degradation at scale.

Skills

Required

  • AI product strategy
  • machine learning
  • AI engineering
  • Python engineering
  • agentic frameworks (LangChain, LangGraph, AutoGen, Google ADK, or equivalent)
  • designing and operating evaluation infrastructure for AI systems
  • Gen AI ecosystem understanding
  • communicating complex AI system design

Nice to have

  • Model Context Protocol (MCP)
  • inference optimization
  • responsible AI
  • full LLM fine-tuning lifecycle
  • published research

What the JD emphasized

  • applying AI at the scale and quality consumers expect
  • drive how AI creates real benefit and support across the organization
  • bring state-of-the-art engineering hands-on execution to building agentic systems
  • Prototyping to Shipping
  • take features from concept to production without a handoff layer
  • production-grade agentic systems
  • Build reusable agent primitives and infrastructure
  • Define AI evaluation as a first-class discipline
  • scale core capabilities
  • Deep, practical experience building and deploying agentic AI systems in production
  • Proven experience designing and operating evaluation infrastructure for AI systems
  • Comfortable navigating from research prototypes to production-ready systems

Other signals

  • building agentic systems
  • applying AI at scale
  • product strategy and engineering execution