Director, Core Infrastructure Engineering

Oracle Oracle · Enterprise · Seattle, WA +1

Director level role focused on leading a team to design, deploy, operate, and optimize GPU infrastructure for AI/ML customers on Oracle Cloud Infrastructure. Responsibilities include leading engineers, driving automation for cluster deployment, acting as a technical liaison, and optimizing infrastructure performance for demanding AI workloads.

What you'd actually do

  1. Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers.
  2. Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or Oracle Kubernetes Engine (OKE) to streamline operations and enhance productivity.
  3. Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures. Act as a technical liaison between customers, core engineering teams, and support.
  4. Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infra expert role.
  5. Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.

Skills

Required

  • Leadership and people management skills
  • Experience building and managing distributed/cloud software engineering solutions
  • Experience using tools like Ansible, Terraform, Python, containerization technologies (e.g., Docker, Kubernetes) and orchestration tools
  • Solid understanding of networking concepts, security principles, and best practices
  • Excellent problem-solving skills
  • Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization
  • Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala
  • Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads

Nice to have

  • Strong communication and collaboration skills
  • Ability to work effectively in cross-functional teams
  • Ability to convey technical concepts to non-technical stakeholders
  • Develop short, medium, and long-term plans to achieve strategic objectives
  • Interact across functional areas with senior management or executives
  • Influence thinking and gain acceptance from others in sensitive situations

What the JD emphasized

  • AI/ML infrastructure customers
  • GPU and AI/ML environments
  • AI/ML infrastructure
  • advanced GPU solutions
  • AI/ML or HPC workloads

Other signals

  • customer-facing infrastructure support
  • GPU cluster management
  • AI/ML workload optimization
  • automation of infrastructure deployment