Senior/staff Site Reliability Engineer

Pinecone Pinecone · Data AI · Tel Aviv, Israel · R&D

Senior/Staff Site Reliability Engineer for Pinecone, a leading vector database company. The role focuses on ensuring the reliability, deployability, and operability of Pinecone's systems across various enterprise environments (SaaS, BYOC, OEM, air-gapped). Responsibilities include designing deployment patterns, building Kubernetes capabilities, defining upgrade strategies, improving observability, and collaborating with product and customer-facing teams to enhance enterprise readiness. This is a platform-focused SRE role, not traditional operations.

What you'd actually do

  1. Design and improve deployment patterns across SaaS, BYOC, and future enterprise deployment models.
  2. Build Kubernetes-based runtime, deployment, and validation capabilities.
  3. Define safe upgrade, rollback, and version compatibility strategies.
  4. Improve observability, alerting, logging, and customer-visible health signals.
  5. Build tooling for deployment validation, diagnostics, and operational readiness.

Skills

Required

  • Strong Kubernetes experience
  • Experience with at least one major cloud provider: AWS, GCP, or Azure
  • Experience operating production systems with clear reliability ownership
  • Experience with monitoring, alerting, logging, incident response, and operational diagnostics
  • Experience designing or operating SaaS, BYOC, self-managed, or regulated enterprise environments
  • Ability to debug complex systems across application, infrastructure, networking, and customer-environment boundaries
  • Ability to write production-quality code in a language such as Go, Python, Java, Rust, or similar
  • Understanding of CI/CD, release engineering, automated validation, and rollback strategies
  • Clear communication and strong ownership across engineering and customer-facing teams

Nice to have

  • Experience with BYOC, self-managed, OEM, or air-gapped deployments
  • Experience with stateful distributed systems such as databases, search systems, storage systems, streaming systems, or ML infrastructure
  • Experience designing upgrade strategies for distributed or stateful systems
  • Experience with enterprise requirements such as SSO, private networking, audit logs, hardened images, compliance constraints, and customer-controlled infrastructure
  • Experience building deployment tooling, operators, controllers, or internal developer infrastructure
  • Experience turning production incidents or customer escalations into product improvements

What the JD emphasized

  • enterprise environments
  • enterprise reliability
  • enterprise-readiness validation
  • enterprise operational requirements
  • enterprise reliability and deployability
  • enterprise environments
  • enterprise requirements