Production Support Engineer - Hyperforce on Alibaba Cloud

Salesforce Salesforce · Enterprise · Singapore

Salesforce is seeking a Production Support Engineer for their Hyperforce on Alibaba Cloud team. This role focuses on SRE principles, automation, and troubleshooting infrastructure and Java applications to ensure reliable service delivery. Responsibilities include detecting, investigating, and resolving system failures, creating observability tooling, and proactively addressing issues. The role emphasizes proactive automation, improving service design for reliability, and driving self-healing initiatives. The engineer will act as a facilitator between development, infrastructure, and support teams, champion customer advocacy, and ensure operational excellence.

What you'd actually do

  1. detecting, investigating and resolving system failures and complex outages across infrastructure and java applications, including creation of the observability tooling necessary for your success.
  2. monitoring the services, reacting to problems, proactively addressing issues before they affect performance or availability, and working with Engineering teams to define service level objectives and improving service design and implementation to increase reliability through closed-loop feedback.
  3. proactive automation, and targets 50%+ time spent on improving service design for reliability, extending monitoring and operational automation, driving self-healing and resiliency initiatives and game day exercises.
  4. technical troubleshooting and investigation, where engineers serve as experts leading the resolution of complex customer issues and production trends.
  5. bridging gaps by acting as a facilitator between development, infrastructure, and support teams to eliminate finger-pointing during cross-product incidents and drive progress.

Skills

Required

  • BS or MS in Computer Science or a related technical field involving systems engineering
  • 5+ years infrastructure and applications systems engineering experience in enterprise-scale Internet services
  • Experience in analyzing and troubleshooting systems using logging, distributed tracing, stack traces, and debuggers
  • 5+ years experience configuring and managing any of the Public Clouds using CLI/SDKs and automation (Alibaba or AWS preferred)
  • 5+ years experience in at least one of the following languages: Java, Python, Go
  • Experience in Unix/Linux environments with good understanding of operating systems internals
  • Working knowledge of the TCP/IP stack, routing and load balancing technologies
  • Working knowledge of design principles of monitoring and alerting systems
  • Ability to operate in a high-pressure environment, troubleshoot complex issues quickly, and successfully handle multiple priorities
  • Systematic problem-solving approach, coupled with a strong sense of ownership and drive
  • Incident management - Act in key support roles during major incidents e.g. Sev0, Sev1. Also, participate in the technical review of the incident for problem management
  • Experience leading complex technical troubleshooting and investigation of critical customer issues and production trends
  • Proven ability to act as a facilitator between development, infrastructure, and support teams to resolve cross-produ

Nice to have

  • Alibaba Cloud preferred
  • AWS preferred

What the JD emphasized

  • strong automation mindset
  • delivering automation solutions
  • troubleshooting Infrastructure and Java applications
  • proactive automation
  • technical troubleshooting and investigation