Engineering Manager, Platform Reliability

Asana Asana · Enterprise · Warsaw, Poland · Infrastructure Engineering

Engineering Manager for Platform Reliability at Asana, focusing on building and leading a new team to ensure the resilience and scalability of Asana's global platform. The role involves defining reliability roadmaps, establishing operational practices, and embedding reliability into system architecture and delivery, with a strong emphasis on people leadership and technical judgment in backend, infrastructure, or reliability-focused systems.

What you'd actually do

  1. Build and lead a new Platform Reliability team, hiring and developing engineers while setting a clear standard for collaboration, ownership, technical excellence and growth.
  2. Partner with technical leaders to define the roadmap for core reliability systems such as load shedding, rate limiting, circuit breakers, traffic controls, and other platform guardrails.
  3. Establish and evangelize best-in-class operating practices for incident response, postmortems, and proactive risk management through SLOs and capacity planning.
  4. Work across platform and infrastructure teams to embed reliability into architecture, planning, and delivery from the start of new initiatives.
  5. Lead end-to-end projects from scoping and design through rollout, making thoughtful tradeoffs between speed, safety, and long-term maintainability.

Skills

Required

  • 5+ years of engineering management experience
  • backend, infrastructure, or reliability-focused teams experience
  • building and operating services in the cloud (e.g. AWS)
  • using cloud-native systems (e.g. Kubernetes)
  • building and operating production systems at scale
  • understanding of failure modes, resilience, and graceful degradation
  • understanding of distributed systems concepts (backpressure, capacity planning, traffic management, reliability targets)
  • strong software engineering background
  • people leadership
  • technical judgment
  • coaching engineers
  • communication skills
  • collaboration skills

Nice to have

  • curiosity about AI tools and emerging technologies
  • willingness to learn and leverage AI tools

What the JD emphasized

  • core reliability systems
  • platform guardrails
  • incident response
  • postmortems
  • proactive risk management
  • SLOs
  • capacity planning
  • reliability into architecture
  • planning
  • delivery
  • resilience
  • graceful degradation
  • distributed systems concepts
  • backpressure
  • capacity planning
  • traffic management
  • explicit reliability targets