Senior Site Reliability Engineer

Oracle Oracle · Enterprise · Reston, VA +1

Senior Site Reliability Engineer responsible for cloud operations of Oracle National Security Realms, focusing on infrastructure, automation, and ensuring high availability and scalability of Oracle products and services, with specific experience in deploying and maintaining AI infrastructure including clustered GPUs and LLMs.

What you'd actually do

  1. Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence.
  2. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services.
  3. Design and develop designs, architectures, standards, and methods for large-scale distributed systems.
  4. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning
  5. Provide cloud operations for Oracle National Security Realms.

Skills

Required

  • Linux and Unix operating systems
  • Docker, Kubernetes, and Terraform
  • Scripting languages such as Bash, shells, Perl, or Python
  • Proficient with writing services/task automation in any modern development language (e.g. Python, Bash, Ruby, Perl, JavaScript, or Java)
  • Deep knowledge of Linux or Unix OS internals and host-based networking
  • Familiarity with configuration management solutions such as Chef, Puppet, etc
  • Experience with devising, managing, and extending monitoring solutions for large scale environments.
  • Knowledge of cloud computing concepts
  • Experience working in a mission-critical environment (Operations, Technical Support, NOC etc)
  • Proficient with communication skills (writing, organization, learning exchange)
  • Experience executing tasks under change management procedures
  • Experience resolving auto-cut and manual alarms following runbooks
  • A focus on customer satisfaction
  • Specific experience working with deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions

Nice to have

  • Familiarity with core protocols and OSI model (DNS, DHCP, HTTP, TCP/IP)
  • A desire to learn and keep up with modern technologies

What the JD emphasized

  • US Citizenship
  • TS/SCI w/Poly security clearance
  • deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions

Other signals

  • deployment of AI infrastructure
  • clustered GPUs
  • LLM deployment and maintenance
  • model integration