Software Engineer Iii, Observability

Box Box · Enterprise · Warsaw, Poland · Engineering Operations - COGs

Software Engineer III on the Observability Engineering team, responsible for architecting and engineering a mission-critical observability platform for globally distributed services. The role involves building software, frameworks, and tools for reliable operations, managing stability and performance of critical production applications, developing automation for platform reliability, and working with distributed systems like Kubernetes and Kafka. The engineer will also collaborate with other teams, participate in POCs, improve observability systems, and participate in on-call rotations. The company emphasizes an AI-first approach and uses AI-powered development tools.

What you'd actually do

  1. You will build software, frameworks, and tools required for reliable operations of Box's services across multiple cloud environments.
  2. You will manage the stability and operation of several of Box's most critical production applications through application reviews, capacity planning, and performance tuning.
  3. You will be constantly developing automations/tooling for better platform reliability/availability.
  4. You will work with cutting-edge & distributed systems such as Kubernetes, Zookeeper, Istio, Kafka and ElasticSearch.
  5. You will improve our observability as both a developer/maintainer of systems/frameworks

Skills

Required

  • 3+ years of experience using higher level languages (e.g. Java, Go)
  • configuration and maintenance of applications such as web servers, load balancers, databases, storage systems, orchestration and/or observability platforms
  • supporting critical production services
  • hands-on experience with modern cloud technologies like GCP, AWS, or/and Kubernetes

Nice to have

  • experience with service observability best practices and tools (e.g.,OTel, Prometheus or the like)
  • Working knowledge of Terraform, Terragrunt or/and Calico for Kubernetes

What the JD emphasized

  • mission-critical observability platform
  • scale across numerous globally distributed environments
  • leverage best in class cloud technologies and practices
  • modern architecture is designed with scalability and longevity in mind
  • improve our observability as both a developer/maintainer of systems/frameworks