Field Reliability Engineer- Latam

Honeycomb Honeycomb · Enterprise · Brazil · Remote · Customer Success

This role focuses on operating and maintaining customer-facing managed infrastructure, including RaaS and HnyPC deployments. It involves building and maintaining automation for provisioning and managing customer environments, designing monitoring and observability systems, and managing scaling, upgrades, and incident response. The role also serves as a senior technical escalation point for complex customer issues, diagnoses deep infrastructure problems, and partners with customer teams to troubleshoot production issues. Additionally, it contributes to open-source projects like OpenTelemetry, builds reference architectures, and acts as a technical backstop for sales and customer success teams, leading architecture reviews and customer-facing POCs. The role also involves building internal tooling and collaborating with various internal teams.

What you'd actually do

  1. Own and operate customer-facing managed infrastructure including Refinery as a Service (RaaS) and Honeycomb Private Cloud (HnyPC) deployments across multiple AWS accounts and regions.
  2. Build and maintain Terraform modules, Helm charts, and deployment automation for provisioning and managing customer EKS clusters, collector pools, and Refinery instances.
  3. Design and implement monitoring, alerting, and observability for managed service infrastructure - using Honeycomb to monitor Honeycomb.
  4. Manage scaling, upgrades, and incident response for customer deployments, including capacity planning and cost optimization across AWS infrastructure.
  5. Serve as the senior technical escalation point for our most challenging customer situations - production incidents, complex collector configurations, Refinery tuning, and architecture reviews that exceed the scope of standard technical roles.

Skills

Required

  • Kubernetes
  • AWS
  • Terraform
  • Helm
  • OpenTelemetry
  • distributed systems
  • networking
  • incident response
  • observability
  • troubleshooting

Nice to have

  • Refinery tuning
  • collector configurations
  • architecture reviews
  • polyglot service meshes
  • SLO workshops
  • hybrid cloud
  • high-cardinality workloads

What the JD emphasized

  • senior technical escalation point
  • deep infrastructure and observability issues
  • partner directly with customer SRE, platform, and engineering teams
  • on-call rotation for managed services
  • customer-facing POCs and pilots