AI Research Scientist, Post-training - Meta Superintelligence Labs

Meta Meta · Big Tech · Menlo Park, CA

Research Scientist focused on designing, generating, and curating post-training data (SFT, RLHF) to align and enhance frontier AI systems. The role involves researching and optimizing post-training recipes, defining data quality frameworks, and influencing research direction for complex reasoning and agentic tasks across various expert domains. Requires a strong publication record and expertise in data-centric AI and alignment techniques.

What you'd actually do

  1. Design novel methodologies for post-training data collection, curation, and synthetic data generation
  2. Define data quality frameworks and alignment strategies that guide capability development across MSL, particularly for complex reasoning and agentic behaviors
  3. Drive the scientific vision for eliciting high-quality data in expert domains (finance, legal, health, STEM) and complex agentic trajectories (Deep research, computer use, UI generation)
  4. Conduct research to develop and optimize post-training recipes that directly improve model quality
  5. Partner with cross-functional research teams across product and model training to identify and prioritize gaps in model capabilities

Skills

Required

  • Ph.D. in Computer Science, Machine Learning, or a related technical field
  • 3+ years of experience in machine learning research, with a focus on deep learning, data alignment, NLP, or related areas
  • Demonstrated ability to lead technical research projects from conception to production
  • Effective communication skills and experience collaborating with technical leadership
  • Hands-on experience with language model post-training, RLHF, DPO, or related alignment techniques

Nice to have

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Recognized expertise in data-centric AI, post-training methodologies, or complex reasoning data

What the JD emphasized

  • Multiple first-author publications at top-tier peer-reviewed venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or similar) related to language model alignment, synthetic data generation, RLHF, or deep learning
  • Track record of research that has substantially influenced the field of deep learning

Other signals

  • post-training data
  • SFT
  • RLHF
  • frontier AI systems
  • model quality
  • agentic tasks