Human Data Engineer

xAI xAI · AI Frontier · Palo Alto, CA · Human Data

This role focuses on creating and executing human data projects to provide high-quality tasks and evaluations for AI model training. The Human Data Engineer partners with model teams, designs projects to capture signals, drives data yield and evaluation lift, and influences data strategy. They are responsible for end-to-end delivery of data and evaluation projects, building tools to measure data effectiveness, maintaining data integrity, and researching techniques for data collection, annotation, generation, and multimodal integration. The role also involves shaping model behavior through targeted data work and collaborating with operations teams to scale projects.

What you'd actually do

  1. Partner with model and engineering teams to understand needs and translate them into high-value data projects and evaluation strategies.
  2. Own end-to-end delivery of critical data and eval projects that capture meaningful training signals and support rapid model development.
  3. Build tools and systems to measure data effectiveness (data yield, eval lift, usage), and run evaluations yourself to continuously improve quality.
  4. Maintain rigorous data integrity and truthfulness, including validation processes for factual accuracy; prioritize quality over quantity.
  5. Research, implement, and evaluate techniques for data collection, annotation, generation, and multi-modal integration.

Skills

Required

  • Bachelor’s degree in engineering, computer science, or a related STEM discipline, or 4+ years of experience in lieu of a degree
  • Experience collaborating with cross-functional teams (engineering, research, product, or annotation/operations groups)
  • Demonstrated experience analyzing datasets to identify trends, anomalies, quality issues, or integrity problems
  • Experience in a scripting language like Python
  • Experience in SQL or other data analysis tools

Nice to have

  • Master’s degree or higher in a relevant technical field
  • Direct experience curating, evaluating, or improving training or evaluation datasets for large language models or other AI/ML systems
  • Experience designing, supporting, or optimizing annotation tools, labeling interfaces, or data workflows that prioritize factual accuracy and data integrity
  • Familiarity with multimodal data (text + images, code, or other modalities) or domain-specific data (science, mathematics, programming, recent events, etc.)
  • Experience conducting model or dataset evaluations focused on quality, truthfulness, or alignment with product goals

What the JD emphasized

  • high-quality tasks and evaluations
  • high-value Human Data projects
  • meaningful signals
  • data yield and eval lift
  • influence data strategy
  • end-to-end delivery
  • critical data and eval projects
  • rapid model development
  • measure data effectiveness
  • run evaluations yourself
  • continuously improve quality
  • rigorous data integrity and truthfulness
  • validation processes for factual accuracy
  • quality over quantity
  • data collection, annotation, generation
  • multi-modal integration
  • Shape Grok’s behavior and domain performance
  • targeted data work
  • annotation workflows, labeling interfaces, and related tools
  • data quality and integrity
  • scale projects
  • share learnings
  • rapid decisions

Other signals

  • crafting and executing high-value Human Data projects
  • partnering with model teams
  • designing projects that capture meaningful signals
  • driving data yield and eval lift
  • influence data strategy