Manager, Performance Research and Analysis

NVIDIA NVIDIA · Semiconductors · Yokneam, Israel +1

Manager role focused on driving end-to-end performance strategy, characterization, and optimization for NVIDIA AI GPU clusters, specifically for large-scale distributed training and inference workloads. This includes deep evaluation of networking technologies, DPU/storage for AI inference, and strategy for performance observability and dashboards. The role involves root-cause analysis and leading technical teams.

What you'd actually do

  1. Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation NVIDIA AI GPU clusters, focusing on large-scale distributed training and inference workloads.
  2. Deeply evaluate and optimize NVIDIA Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.
  3. Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.
  4. Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.
  5. Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.

Skills

Required

  • B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.
  • 8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.
  • 3+ years of experience as an engineering team manager leading technical performance or R&D teams.
  • Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).
  • Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
  • Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.

Nice to have

  • Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.
  • Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
  • Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.
  • Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.

What the JD emphasized

  • large-scale distributed training and inference workloads
  • AI inference jobs
  • performance observability
  • large-scale distributed training and inference workloads
  • AI inference jobs
  • performance observability

Other signals

  • AI GPU clusters
  • distributed training and inference
  • performance optimization
  • networking technologies
  • performance observability