System Design Engineer – AI Cluster Storage Architect

AMD AMD · Semiconductors · Austin, TX · Engineering

AMD is seeking a System Design Engineer specializing in AI Cluster Storage Architecture. This hands-on role involves researching and developing AI infrastructure, with a primary focus on storage solutions for AI/HPC clustered systems. The engineer will create reference architectures, configuration guides, and conduct benchmarking experiments to support internal teams and customers. The role requires deep technical evaluation of HPC and AI stacks, particularly storage, and developing technical artifacts to enable others.

What you'd actually do

  1. Apply your HPC expertise to shape AI infrastructure by creating reference architectures, configuration guides, and deployment blueprints that help internal teams and customers make informed hardware and software decisions
  2. Build a library of technical artifacts—including presentations, design documents, and “how it works” guides, to support pre-sales engineers and enable others to skill up from an HPC perspective
  3. Perform deep technical evaluations of HPC and AI stacks with a specific focus on storage solutions, documenting how they work, where they fit, and the tradeoffs involved between storage technologies and vendors
  4. Design and execute reproducible experiments and benchmarking harnesses to compare storage technologies and their fit-for-purpose across AI and HPC workloads
  5. Develop small reference implementations and tools to validate performance hypotheses, analyze system behavior and more

Skills

Required

  • HPC expertise
  • systems thinking
  • troubleshooting
  • Linux fundamentals
  • Clear communication
  • writing technical artifacts

Nice to have

  • parallel filesystems (Lustre, BeeGFS)
  • object stores
  • RDMA
  • data pipeline throughput and caching strategies
  • storage vendor landscape
  • schedulers and/or orchestration systems (e.g., Slurm, Kubernetes)
  • MPI/OpenMP
  • distributed storage patterns
  • evaluation docs/RFCs
  • networking
  • containers
  • performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF)
  • AMD ecosystem experience (ROCm, RCCL, Instinct GPUs, EPYC platforms)
  • Distributed training internals (DDP, collective comms, sharded/stateful optimizers; NCCL/RCCL behavior)
  • Orchestration models (Slurm configuration patterns, Kubernetes for HPC/AI, Apptainer/Singularity)
  • Enterprise storage solutions (NAS, NFS)
  • IaC literacy (Terraform/Ansible)

What the JD emphasized

  • hands-on
  • researching and experimentation
  • storage solutions for AI/HPC
  • benchmarking
  • reference architecture
  • reproducible experiments
  • deep technical evaluations
  • comparative analysis
  • performance analysis

Other signals

  • AI infrastructure
  • storage solutions for AI/HPC
  • benchmarking
  • reference architectures
  • reproducible experiments