Software Engineer - Platform Infrastructure (rust, C++)

xAI xAI · AI Frontier · Palo Alto, CA · Engineering

Software Engineer focused on building and optimizing large-scale distributed systems and infrastructure for AI training clusters, involving low-level systems programming and hardware/software co-design.

What you'd actually do

  1. Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
  2. Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  3. Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  4. Maintain and innovate on our codebase to ensure scalability and reliability.
  5. Develop tools to enhance team productivity and streamline workflows.

Skills

Required

  • Systems programming experience in C, C++, or Rust
  • Computer systems fundamentals
  • Hands-on expertise with Kubernetes (K8s)

Nice to have

  • Strong debugging skills across the full stack
  • Deep knowledge of operating systems internals
  • Proficiency in performance analysis, profiling, and low-level optimization techniques
  • Solid understanding of computer networks and the TCP/IP stack
  • Experience working with Linux kernel concepts or systems-level debugging tools
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
  • Solid understanding of containerization technologies
  • Experience with observability and monitoring in distributed systems

What the JD emphasized

  • low-level stack
  • AI training
  • hardware, software, and algorithm co-design

Other signals

  • large-scale distributed system
  • supercomputing clusters
  • AI training
  • low-level stack optimization
  • hardware, software, and algorithm co-design