Staff Engineer – Observability Platform

Postman Postman · Enterprise · Bangalore, India · Platform Engineering

Staff Engineer role focused on building and driving the technical vision and architecture for Postman's observability platform, including metrics, logging, tracing, and alerting. The role involves investigating production issues, establishing observability standards, and building tooling for faster incident detection and resolution. Requires strong experience in distributed systems, cloud-native architectures, and observability technologies.

What you'd actually do

  1. Drive the technical vision and architecture for Postman's observability platform.
  2. Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics.
  3. Investigate complex production issues, identify root causes, and drive long-term corrective actions.
  4. Partner with engineering teams to improve service reliability, availability, performance, and operational maturity.
  5. Establish observability standards, best practices, and instrumentation frameworks across the company.

Skills

Required

  • 10+ years of software engineering experience
  • distributed systems
  • cloud-native architectures
  • observability domains (monitoring, logging, distributed tracing, telemetry pipelines, incident management)
  • operating large-scale production systems
  • reliability
  • scalability
  • performance
  • system debugging
  • root-cause analysis
  • performance optimization
  • production operations
  • Go, Java, Python, Node.js, or similar programming languages
  • OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb, or equivalent platforms
  • influence technical direction
  • communication skills
  • data-driven approach

Nice to have

  • building internal developer platforms or observability platforms at scale
  • AIOps, intelligent alerting, anomaly detection, or AI-powered operational tooling
  • driving reliability initiatives across multiple engineering organizations
  • SRE, Platform Engineering, Infrastructure Engineering, or Developer Productivity background

What the JD emphasized

  • 10+ years of software engineering experience
  • significant exposure to distributed systems
  • cloud-native architectures
  • Strong expertise in observability domains
  • operating large-scale production systems
  • strong focus on reliability
  • scalability
  • performance
  • Deep understanding of system debugging
  • root-cause analysis
  • performance optimization
  • production operations