Senior Software Engineer, Observability

Together AI Together AI · Data AI · San Francisco, CA · Engineering

Senior Software Engineer focused on building and scaling a robust observability platform for AI infrastructure, including metrics, logs, traces, monitoring, alerting, and anomaly detection. The role involves designing and implementing scalable systems using tools like Prometheus, Grafana, ClickHouse, OpenTelemetry, Go, Python, and Terraform, with a focus on distributed systems, containerization, and orchestration.

What you'd actually do

  1. Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows.
  2. Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
  3. Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
  4. Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
  5. Define observability best practices.

Skills

Required

  • Prometheus
  • Grafana
  • ClickHouse
  • ClickStack
  • OpenTelemetry
  • Go
  • Python
  • Terraform
  • Ansible
  • Helm
  • Docker
  • Kubernetes
  • PostgreSQL
  • MongoDB
  • Redis
  • time-series databases

Nice to have

  • AI/ML infrastructure monitoring
  • GPU cluster monitoring
  • model performance metrics
  • training pipeline monitoring
  • high-frequency systems monitoring
  • low-latency systems monitoring
  • chaos engineering
  • reliability testing
  • open-source observability projects
  • security monitoring
  • compliance frameworks

What the JD emphasized

  • scalable observability platform
  • metrics, logs, traces
  • Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry
  • telemetry data pipelines
  • log aggregation workflows
  • automated monitoring, alerting, and anomaly detection systems
  • SLIs/SLOs
  • runbooks
  • predictive analytics
  • critical services
  • custom observability tools
  • infrastructure-as-code
  • Go, Python, Terraform, Ansible, and Helm
  • distributed tracing
  • application monitoring
  • incident response
  • post-mortem analysis
  • observability best practices
  • Expertise in observability platforms
  • Prometheus, Grafana, ClickStack, OpenTelemetry
  • cloud-native monitoring services
  • Strong programming skills in Go, Python, or similar languages
  • proficiency in infrastructure-as-code tools
  • Terraform, Ansible, Helm
  • Experience designing, operating, and scaling large-scale distributed systems and pipelines
  • high-volume data ingestion
  • real-time querying
  • Deep understanding of containerization
  • orchestration
  • Kubernetes
  • microservices architecture
  • service mesh technologies
  • CI/CD pipelines
  • GitOps workflows
  • Expertise in managing databases
  • PostgreSQL, MongoDB, Redis
  • time-series databases
  • high-cardinality data