Senior Machine Learning Engineer

Zendesk Zendesk · Enterprise · Melbourne, Australia +1 · Remote

Zendesk is seeking a Senior Machine Learning Engineer to architect and build the core infrastructure for their AI Agent runtime. This role focuses on creating a fault-tolerant, multi-tenant execution environment with robust state management, memory functions, resource isolation, and integration capabilities for AI agents. The goal is to enable LLMs to act as reliable digital employees within their CX platform, driving AI-driven revenue.

What you'd actually do

  1. You will move beyond basic loops to architect fault-tolerant execution environments based on the principles of durable execution: incremental execution, state persistence, and fault tolerance via automated recovery.
  2. You will design convergent, high-throughput memory infrastructures to support the working, episodic, semantic, and procedural memory tiers of our agents.
  3. You will ensure strict logical boundaries across shared infrastructure pools.
  4. You will drive our adoption of the Model Context Protocol (MCP), building the secure servers and clients that allow agents to dynamically discover tools, negotiate capabilities, and lay the groundwork for secure Agent-to-Agent (A2A) collaboration.
  5. You will establish the telemetry required to trace agent actions back to specific tenants, enforcing rate limits and optimizing the compute cost-per-token to ensure the unit economics of the platform remain highly viable.

Skills

Required

  • Java, Go, or Python backend experience
  • Distributed state management
  • Event-driven architectures (Kafka)
  • Non-deterministic execution in cloud environments
  • Durable execution frameworks
  • Event-sourcing orchestration
  • High-throughput in-memory datastores
  • Converged databases
  • Vector stores
  • Postgres
  • Kafka
  • Model Context Protocol (MCP)
  • gRPC
  • GraphQL
  • Systemic architectural bottleneck identification
  • Technical direction definition for platforms
  • Mentoring engineering teams on agentic AI

Nice to have

  • Advanced durable execution frameworks
  • Advanced event-sourcing orchestration

What the JD emphasized

  • strictly dedicated to infrastructure
  • multi-tenant scaling
  • architect fault-tolerant execution environments
  • state persistence
  • fault tolerance
  • automated recovery
  • long-running workflows
  • state management systems
  • strict ACID transactions
  • strict logical boundaries
  • zero-trust execution validation
  • native Row-Level Security
  • secure servers and clients
  • secure Agent-to-Agent (A2A) collaboration
  • trace agent actions back to specific tenants
  • rate limits
  • optimize the compute cost-per-token
  • unit economics
  • extensive backend experience
  • deeply understand the nuances of managing distributed state
  • event-driven architectures
  • non-deterministic execution
  • reliability in an AI system
  • externalizing state
  • safe self-correction
  • deterministic recovery
  • extreme complexities of API design
  • capability negotiation
  • secure sandboxing for tool execution

Other signals

  • building agentic systems
  • multi-tenant runtime
  • distributed state management
  • enterprise memory functions
  • integration fabric for agents