ML Infrastructure Engineer

xAI xAI · AI Frontier · Palo Alto, CA · Engineering

xAI is seeking an ML Infrastructure Engineer to build and optimize the reliable, high-performance ML platform powering their recommendations. This role involves designing and scaling GPU compute infrastructure, training frameworks, and experimentation tools, developing data pipelines, and integrating large-scale data, training, and inference systems. The engineer will collaborate with ML teams to productionize models, ensure scalability, reliability, and efficiency, and work across the full stack. Requires 2+ years of industry experience with large-scale production environments, distributed systems, GPU infrastructure, and ML platforms, proficiency in Python and C++/Rust, and familiarity with ML frameworks like JAX/PyTorch and Linux/orchestration tools.

What you'd actually do

  1. Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
  2. Developing data pipelines and integrating large-scale data, training, and inference systems
  3. Collaborating with ML teams to productionize models and ensure seamless integration across the stack
  4. Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
  5. Working across the full stack to solve complex problems independently

Skills

Required

  • 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
  • 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
  • Strong proficiency with Python
  • experience with compiled languages such as C++ or Rust

Nice to have

  • Deep familiarity with modern ML frameworks such as JAX or PyTorch
  • Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
  • Comfortable with Linux systems and orchestration tools
  • Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling

What the JD emphasized

  • high-performance ML platform
  • large-scale data
  • large-scale machine learning systems

Other signals

  • GPU compute infrastructure
  • training frameworks
  • experimentation tools
  • data pipelines
  • large-scale data
  • training systems
  • inference systems
  • productionize models
  • scalability
  • reliability
  • efficiency
  • large-scale machine learning systems
  • full stack
  • Python
  • C++
  • Rust
  • JAX
  • PyTorch
  • Linux
  • job schedulers
  • configuration management