Applied AI Engineer, Inference

Weights & Biases Weights & Biases · Data AI · Bellevue, WA +3 · Remote · Technology

Applied AI Engineer focused on optimizing and benchmarking the inference serving platform for AI models, working on latency, throughput, and reliability.

What you'd actually do

  1. Build and maintain benchmarking workflows that measure latency, throughput, quality regressions, and cost across priority models and serving configurations.
  2. Benchmark our inference stack against realistic customer workloads and external provider baselines to identify performance gaps and improvement opportunities.
  3. Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  4. Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations, evaluating runtime features, and validating changes against representative traces and benchmarks.
  5. Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching strategies, routing, and other inference optimizations, with careful attention to quality and correctness tradeoffs.

Skills

Required

  • Python
  • LLM inference systems
  • vLLM
  • SGLang
  • TensorRT-LLM
  • model-serving stacks
  • latency
  • throughput
  • batching
  • GPU utilization
  • quantization
  • quality regression analysis
  • empirical evaluations
  • benchmarks
  • experiments
  • production engineering environments

Nice to have

  • optimizing inference workloads on modern GPU hardware
  • Nsight Systems
  • PyTorch profilers
  • telemetry pipelines
  • benchmark suites
  • evaluation frameworks for coding, reasoning, or agent workloads
  • real production traces
  • customer traffic patterns
  • speculative decoding
  • acceleration strategies

What the JD emphasized

  • real-world performance
  • rigorous benchmarks
  • performance bottlenecks
  • targeted optimizations
  • applied performance work
  • performance testing
  • performance goals
  • performance engineering
  • performance engineering
  • performance engineering
  • performance improvements

Other signals

  • performance optimization
  • benchmarking
  • inference serving