Senior Performance Engineer

NVIDIA NVIDIA · Semiconductors · Yokneam, Israel

Senior Performance Engineer at NVIDIA focusing on optimizing AI and HPC workloads on GPU/CPU clusters. Responsibilities include profiling, benchmarking, identifying bottlenecks, and developing performance analysis tools. Requires 5+ years of experience in performance analysis, systems engineering, or HPC/AI infrastructure, with expertise in high-performance networking and system performance metrics.

What you'd actually do

  1. Profile, benchmark, and analyze AI and HPC workloads on GPU and CPU clusters
  2. Explore performance characteristics of high-performance networking and collective communications (e.g., NCCL, RDMA, MPI, RoCE)
  3. Identify performance bottlenecks across networking, compute, memory, and system architecture
  4. Develop and enhance performance analysis, benchmarking, and diagnostic tools
  5. Define performance test plans and establish expectations for new technologies and platforms

Skills

Required

  • performance analysis
  • systems engineering
  • HPC/AI infrastructure
  • high-performance networking
  • system performance metrics
  • telemetry environments
  • analytical skills
  • problem-solving skills
  • communication skills
  • cross-functional collaboration

Nice to have

  • CUDA
  • NCCL internals
  • congestion control algorithms
  • CPU architectures
  • GPUs
  • HCAs
  • memory
  • PCIe
  • NVIDIA GPUs
  • deep learning frameworks
  • PyTorch
  • TensorFlow
  • cloud platforms
  • Python
  • Bash
  • C/C++
  • Linux environments

What the JD emphasized

  • performance analysis
  • AI and HPC workloads
  • high-performance networking
  • system performance metrics

Other signals

  • performance analysis
  • GPU and CPU clusters
  • AI and HPC workloads
  • high-performance networking