Staff Platform Engineer, Service Infrastructure

Together AI Together AI · Data AI · San Francisco, CA · Engineering

Staff Platform Engineer focused on service infrastructure strategy, including Kubernetes, AWS, Terraform, networking, and observability. The role involves improving reliability, operability, and production readiness of core platforms, and driving cross-company infrastructure initiatives.

What you'd actually do

  1. Own the technical direction for service infrastructure within Product Foundations, including Kubernetes, AWS, Terraform, CDNs, ALBs, DNS, IAM, service networking, and related operational patterns.
  2. Up-level existing Product Foundations services by improving reliability, operability, deployment safety, infrastructure consistency, and production readiness.
  3. Partner deeply with API Platform and UI Platform on networking, DNS, CDN, load balancing, delivery, and gateway patterns for critical customer-facing interfaces.
  4. Work closely with Infrastructure, Networking, and Security teams to bring company-wide platform standards into Product Foundations and contribute PF requirements back into shared frameworks.
  5. Help drive cross-company infrastructure initiatives that Product Foundations depend on or help maintain, including Terraform CI/CD, Kubernetes networking, zero-trust service communication, policy-as-code, and cross-DC/provider networking.

Skills

Required

  • Kubernetes
  • AWS
  • Terraform
  • CDNs
  • ALBs
  • DNS
  • IAM
  • service networking
  • Helm
  • ArgoCD/Argo Rollouts
  • ingress
  • autoscaling
  • secrets
  • service identity
  • progressive delivery
  • module design
  • infrastructure CI/CD
  • policy enforcement
  • production applies
  • safe self-service workflows
  • operating networking and edge infrastructure
  • TLS
  • traffic management
  • Go
  • Python
  • TypeScript
  • EKS
  • VPC networking
  • Route 53
  • CloudFront
  • ECR
  • observability systems
  • metrics
  • logs
  • traces
  • dashboards
  • alerting
  • SLOs
  • incident response
  • lead cross-functional technical initiatives
  • design docs
  • architecture reviews
  • mentorship
  • hands-on implementation
  • pragmatic tradeoffs
  • influence without authority

Nice to have

  • internal developer platforms
  • paved-path service frameworks
  • service mesh
  • zero-trust infrastructure
  • mTLS
  • SPIFFE/SPIRE
  • Cilium
  • Istio
  • Linkerd
  • Envoy
  • OPA
  • Gatekeeper
  • Kyverno
  • Sentinel
  • policy-as-code
  • multi-region
  • multi-cluster
  • hybrid-cloud
  • cross-provider service networking
  • supply-chain security
  • image signing
  • SBOMs
  • vulnerability management
  • compliance automation

What the JD emphasized

  • 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles.
  • Deep production experience with Kubernetes, including EKS, Helm, ArgoCD/Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery.
  • Strong Terraform experience, including module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows.
  • Staff-level judgment: you can define ambiguous problems, make pragmatic tradeoffs, influence without authority, and leave both systems and teams better than you found them.