Key Responsibilities:
- Deploy, monitor, maintain, and optimize production server hardware, software, and databases within a cloud environment.
- Deliver escalated technical support for complex and critical issues, engaging in problem management and communicating status updates to stakeholders.
- Lead resolution of escalated support cases, collaborating with internal technical resources and third-party vendors when necessary.
- Manage storage infrastructure and coordinate upgrades, bug fixes, and patching for Oracle systems and database appliances.
- Support standardization and automation initiatives across hardware and software within the Oracle technology stack to enhance reliability and supportability.
- Maintain exceptional system uptime and ensure systems meet the rigorous performance and availability expectations of cloud-native environments.
- Participate in an on-call rotation to provide after-hours support for production systems and respond to urgent incidents.
Preferred Qualifications:
- Strong experience in administration and analysis of cloud-based production environments, preferably within Oracle Cloud Infrastructure.
- Advanced experience with Linux systems administration.
- Demonstrated expertise in troubleshooting complex technical problems related to scalability and high availability.
- Proficiency in managing server operating systems, storage environments, and database appliances.
- Effective communicator with strong problem-solving skills and the ability to lead cross-functional resolution efforts.
- Familiarity with automation tools, standardization projects, and best practices for maintaining production environments in the cloud.
- Willingness to participate in an on-call rotation as required.
- This position is a hybrid role which does require employees to be in office 3 day a week at our Guadalajara office.
Responsibilities:
Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.
Career Level - IC3