Senior Machine Learning Engineer

Zendesk Zendesk · Enterprise · Melbourne, Australia +1 · Remote

Zendesk is seeking a Senior Machine Learning Engineer to advance their AI agent platform, focusing on agent orchestration, planning, multi-agent coordination, and autonomous task execution for customer service resolution. The role involves pushing architectural boundaries, developing domain-specialized agent models through RL, hardening evaluation systems, and implementing enterprise-scale guardrails. Experience with production ML/AI systems, agent architectures, and evaluation is required, with a preference for RL expertise.

What you'd actually do

  1. Pushing the architecture further.
  2. Domain-specialized agent models.
  3. Hardening evaluation.
  4. Guardrails at enterprise scale.

Skills

Required

  • production ML/AI systems
  • agent architectures
  • planning
  • tool dispatch
  • memory
  • error recovery
  • evaluation
  • Python
  • PyTorch

Nice to have

  • RL for language models
  • reward shaping
  • online/offline tradeoffs
  • reward hacking
  • agent framework

What the JD emphasized

  • 5+ years building production ML/AI systems, with hands-on experience in agent architectures (planning, tool dispatch, memory, error recovery). If you have only used LangChain tutorials, this is not the right fit.
  • Strong evaluation instincts. You understand why public benchmarks diverge from production performance and you have built internal evals to close that gap.
  • Python and PyTorch fluency.
  • multi-agent coordination
  • autonomous learning mechanisms
  • autonomous resolve customer service tickets
  • without a human in the loop
  • self-learning mechanism
  • improves from its own execution history
  • multi-step tool-use benchmarks
  • internal evaluation suite
  • scenario-based tests
  • regression detection
  • pushing the architecture further
  • plan decomposition
  • memory tiers
  • skill acquisition
  • multi-agent delegation
  • specialized agents
  • Agent-to-Agent protocol
  • domain-specialized agent models
  • training our own models
  • specialized for customer service resolution
  • RL on production trajectories
  • data pipeline
  • RL training infrastructure
  • reward curricula
  • rollout systems
  • feedback loops
  • hardened evaluation
  • multi-turn evaluation
  • automated trajectory analysis
  • quality gates
  • CI integration
  • guardrails at enterprise scale
  • tool misuse
  • cascading action chains
  • prompt injection
  • hallucination loops
  • multi-layered defenses
  • supervisor patterns
  • capabilities-based access control
  • output validation
  • thousands of concurrent sessions
  • meaningful latency

Other signals

  • building production AI agents
  • multi-step resolution
  • autonomous resolution
  • real APIs
  • without a human in the loop
  • iterative architecture
  • self-learning mechanism
  • synthesized into new skills
  • improves from its own execution history
  • multi-step tool-use benchmarks
  • internal evaluation suite
  • scenario-based tests
  • regression detection
  • pushing the architecture further
  • plan decomposition
  • memory tiers
  • skill acquisition
  • multi-agent delegation
  • specialized agents
  • Agent-to-Agent protocol
  • domain-specialized agent models
  • training our own models
  • specialized for customer service resolution
  • RL on production trajectories
  • data pipeline
  • RL training infrastructure
  • reward curricula
  • rollout systems
  • feedback loops
  • hardened evaluation
  • multi-turn evaluation
  • automated trajectory analysis
  • quality gates
  • CI integration
  • guardrails at enterprise scale
  • tool misuse
  • cascading action chains
  • prompt injection
  • hallucination loops
  • multi-layered defenses
  • supervisor patterns
  • capabilities-based access control
  • output validation
  • thousands of concurrent sessions
  • meaningful latency