Senior Software Engineer - Perf and Benchmarking

Weights & Biases Weights & Biases · Data AI · Bellevue, WA +1 · Technology

Senior Software Engineer focused on performance and benchmarking for AI infrastructure, specifically developing and enhancing Kubernetes-native services for measuring latency, throughput, and cost. The role involves contributing to MLPerf Training and Inference runs, optimizing performance across the compute stack, and working with distributed systems and high-performance computing components, including GPU systems and model-serving stacks.

What you'd actually do

  1. Develop and enhance Kubernetes-native benchmarking services that measure latency, throughput, jitter, and cost-per-request across CoreWeave’s compute stack.
  2. Contribute to implementing and maintaining benchmarking workflows for end-to-end MLPerf Training and Inference runs, including workload setup, cluster configuration, and result validation.
  3. Participate in design discussions and contribute to architecture decisions within the team.
  4. Break down engineering tasks into clear milestones and deliver reliable, high-quality code.
  5. Collaborate with teammates to maintain reproducible, well-documented benchmarking processes.

Skills

Required

  • Python or Go
  • Kubernetes
  • distributed systems
  • high-performance computing
  • cloud services
  • networked systems
  • performance fundamentals
  • CI/CD
  • observability tools

Nice to have

  • time-series databases
  • LSM-based storage engines
  • custom data pipelines
  • MLPerf
  • large-scale benchmarking frameworks
  • llm-d
  • vLLM
  • PyTorch
  • benchmarking GPU clusters
  • multi-region environments
  • CUDA kernels
  • NCCL/SHARP
  • RDMA/NUMA
  • GPU interconnect topologies

What the JD emphasized

  • industry-leading end-to-end performance benchmarking publications such as MLPerf
  • strict P99 SLAs at scale
  • performance-critical GPU systems
  • model-serving stacks

Other signals

  • performance benchmarking
  • MLPerf
  • Kubernetes-native benchmarking services
  • GPU systems
  • model-serving stacks