Data Scientist, Inference Capacity Optimization

OpenAI OpenAI · AI Frontier · San Francisco, CA · Scaling

Data Scientist role focused on optimizing inference capacity for OpenAI's global GPU fleet by building statistical and ML models for utilization, latency, and demand forecasting. This role supports infrastructure planning and investment decisions for AI compute environments.

What you'd actually do

  1. Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency.
  2. Develop forecasting models for inference demand across products, regions, and model families.
  3. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities.
  4. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies.
  5. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs.

Skills

Required

  • Python
  • SQL
  • forecasting
  • optimization
  • predictive models
  • experimentation
  • statistical inference
  • causal analysis
  • communicate analytical insights to executive stakeholders

Nice to have

  • Capacity planning
  • Distributed systems
  • AI infrastructure
  • Datacenter design and buildout
  • Queueing theory
  • Time-series forecasting
  • Operations research
  • Supply-demand modeling
  • Reinforcement learning for resource allocation
  • Cost optimization

What the JD emphasized

  • 5+ years of experience working in the infrastructure data science space

Other signals

  • optimize inference capacity
  • GPU utilization
  • latency, throughput
  • forecasting models for inference demand
  • identify latency bottlenecks and capacity constraints