Staff Software Engineer, Graphql

Airbnb Airbnb · Consumer · United States · Software Engineering

Staff Software Engineer role focused on the Viaduct GraphQL platform, driving reliability, operational excellence, and developer experience. The role involves designing and implementing deployment pipelines, SLO frameworks, observability tooling, and AI-powered operational tooling for incident triage and debugging. It also emphasizes an AI-first engineering approach using LLM-powered agents for code generation and iteration, and contributing to the open-source Viaduct project.

What you'd actually do

  1. Drive platform reliability and operational excellence by designing and implementing deployment pipelines, SLO frameworks, observability tooling, performance improvements, and AI-enabled incident response automation that help maintain Viaduct's 99.99% uptime target across Airbnb's critical API traffic.
  2. Architect and deliver AI-powered operational tooling that accelerates incident triage, reduces mean-time-to-mitigation, and empowers both the Viaduct team and tenant engineers with self-service debugging capabilities.
  3. Embrace an AI-first engineering approach, using LLM-powered agents to generate and iterate on code while you focus on problem-solving, system design, and quality oversight.
  4. Shape the future of Viaduct Modern by contributing to the next-generation architecture, improving developer experience for hundreds of engineers, and establishing patterns that will be shared with the open-source community.
  5. Contribute to open-source Viaduct by ensuring platform improvements are generalizable and well-documented for the broader engineering community.

Skills

Required

  • 9+ years of software engineering experience
  • backend systems
  • distributed architectures
  • platform engineering
  • observability and monitoring
  • SLO frameworks
  • distributed tracing systems
  • metrics pipelines
  • reliability engineering
  • incident response
  • root cause analysis
  • high availability systems
  • performance tuning
  • resource management
  • JVM-based systems
  • profiling
  • garbage collection optimization
  • concurrency models
  • deployment safety
  • automated rollbacks
  • progressive delivery strategies
  • GraphQL
  • API gateway/data access layer technologies
  • developer tooling and platforms
  • product mindset
  • leadership and communication skills

Nice to have

  • Kotlin
  • coroutines

What the JD emphasized

  • AI-enabled incident response automation
  • AI-powered operational tooling
  • LLM-powered agents
  • AI-first engineering approach
  • 99.99% uptime target

Other signals

  • AI-enabled incident response automation
  • AI-powered operational tooling
  • LLM-powered agents
  • AI-first engineering approach