[2026] Data Scientist, Foundation AI - Phd Early Career

Roblox Roblox · Consumer · San Mateo, CA · Early Career Full-Time

Roblox is seeking a Data Scientist for their Foundation AI team, focusing on evaluating and optimizing user-facing Generative AI systems. The role involves developing evaluation frameworks, running experiments, defining success metrics, building automated systems, and conducting applied research. The ideal candidate will have a PhD, strong technical skills in SQL and Python, expertise in experimentation and causal inference, and familiarity with GenAI models and evaluation methods. Experience with model training lifecycle (fine-tuning, RLHF, synthetic data) and a track record of applied research or publications are highly valued. The role is primarily focused on the evaluation gate (L5) for GenAI features, with a secondary focus on fine-tuning and RLHF methods (L2).

What you'd actually do

  1. Design and operationalize rigorous evaluation systems for either GenAI features (text, image, video, 3D, 4D). This includes eval experiment design, dataset design, label reliability analysis, and implementing and finetuning LLM-as-judge methods.
  2. Conduct online experiments (A/B tests) and causal inference to quantify the impact of GenAI features. You will identify opportunities, measure lift, and ensure statistical rigor.
  3. Partner with cross-functional teams to define leading/lagging indicators for GenAI feature user satisfaction, business success, and safety.
  4. Research and apply state-of-the-art methodologies to build reproducible evaluation tooling that lift rigor and efficiency across the company.
  5. Maintain an active pulse on the intersection of Gen AI and Data Science. You will innovate on methodology and techniques to solve unique business challenges while contributing to the broader field in the technical community.

Skills

Required

  • PhD or pursuing PhD in a quantitative field
  • SQL (Hive/Spark)
  • Python or R
  • Experimentation and Causal Inference
  • Statistical Analysis
  • Problem Framing and Solving
  • GenAI Familiarity

Nice to have

  • Expertise in the model training lifecycle (e.g., fine-tuning, RLHF, or synthetic data generation)
  • Applied Research Background or publications

What the JD emphasized

  • PhD or equivalent
  • rigorous evaluation systems
  • LLM-as-judge methods
  • GenAI Familiarity
  • model training lifecycle is a plus
  • applied research or publications

Other signals

  • evaluation frameworks
  • GenAI features
  • safety, responsibility, quality, and efficiency
  • LLM-as-a-judge
  • applied research