Principal Core Infrastructure Engineer

Oracle Oracle · Enterprise · Seattle, WA +1

Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer) at Oracle. This role focuses on designing, implementing, and maintaining the infrastructure that supports customer AI and machine learning initiatives, working closely with customer technical teams and internal cloud services to ensure efficient, secure, and scalable AI/ML solutions. Responsibilities include optimizing infrastructure, troubleshooting POC and production deployments, collaborating with scientists and engineers on infrastructure requirements for ML models, implementing automation, optimizing performance, ensuring security and compliance, and acting as a technical liaison.

What you'd actually do

  1. Design, deploy, and manage infrastructure components such as cloud resources, distributed computing systems, and data storage solutions to support AI/ML workflows.
  2. Collaborate with scientists and software/infrastructure engineers to understand infrastructure requirements for training, testing, and deploying machine learning models.
  3. Implement automation solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity.
  4. Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.
  5. Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.

Skills

Required

  • scripting and automation using tools like Ansible, Terraform, Python and/or Kubernetes
  • containerization technologies (e.g., Docker, Kubernetes)
  • orchestration tools (like Slurm, PBS, etc.)
  • networking concepts
  • security principles
  • problem-solving skills
  • communication and collaboration skills
  • documentation skills
  • Linux skills
  • system administration
  • package management
  • shell scripting
  • performance optimization

Nice to have

  • Python, Rust, Go, Java, or Scala
  • designing, implementing, and managing infrastructure for AI/ML or HPC workloads
  • machine learning frameworks and libraries such as TensorFlow, PyTorch, or sci-kit-learn
  • DevOps practices and tools
  • High-Performance Computing/GPU systems

What the JD emphasized

  • AI/ML Forward Deployed Infrastructure Engineer
  • AI and machine learning initiatives
  • AI/ML solutions
  • AI/ML infrastructure
  • training, testing, and deploying machine learning models
  • AI/ML infrastructure stack
  • AI/ML or HPC workloads

Other signals

  • customer-facing role
  • infrastructure for AI/ML
  • deploying ML models
  • optimizing infrastructure
  • performance, reliability, cost-effectiveness
  • POC and production deployments
  • GPU solutions