AI Tutor - Catalan

xAI xAI · AI Frontier · United States · Remote · Human Data

The role involves training and refining AI models (Grok) for multilingual audio capabilities, focusing on voice interactions, speech recognition, and auditory experiences. Responsibilities include annotating audio data, translating and localizing text, and collaborating with technical staff to improve AI's handling of diverse audio nuances and annotation tools. The core work is data curation and annotation for AI model training.

What you'd actually do

  1. Use proprietary software to provide labels, annotations, recordings, and inputs on projects involving multilingual audio clips, voice recordings, speech samples, and auditory elements in various languages.
  2. Translate and localize text strings (UI copy, prompts, responses, and other written content) from English into Catalan, ensuring linguistic accuracy, natural phrasing, and cultural appropriateness.
  3. Support the delivery of high-quality curated audio data that ensures clear, natural spoken output, accurate representation of linguistic and prosodic details (such as intonation, rhythm, and accent), and professional audio standards.
  4. Collaborate with technical staff to develop tasks that improve AI's ability to handle speech modulation, accent variation, noise in real-world recordings, and multilingual audio processing.
  5. Work with technical staff to improve annotation tools for efficient audio workflows.

Skills

Required

  • Native proficiency in Catalan
  • Proficiency in English (minimum B2 level)
  • Clear, natural vocal delivery and pronunciation
  • Strong auditory perception
  • Ability to handle multilingual audio content
  • Ability to transcribe audio with high accuracy
  • Ability to translate and localize text accurately between English and Catalan
  • Comfort providing high-quality voice recordings
  • Strong comprehension skills
  • Ability to make independent judgments on ambiguous or varied audio material
  • Strong communication skills
  • Strong interpersonal skills
  • Strong analytical skills
  • Detail-oriented
  • Organizational skills
  • Ability to articulate audio-related feedback effectively

Nice to have

  • Exceptional attention to linguistic nuance, auditory detail, and data quality
  • Deep understanding and taste of what good/useful Audio data is
  • Strong command of advanced transcription and annotation practices
  • Background in linguistics (e.g., phonetics, phonology, sociolinguistics), speech sciences, cognitive science, or a related field
  • Experience working with speech/audio datasets, annotation workflows, or AI training data
  • Knowledge/experience with training voice models
  • Understanding of how data quality impacts model performance
  • Experience with translation or localization workflows
  • Professional experience in voice work, including voice acting, voice recording, podcasting
  • Portfolio (strongly preferred for advanced candidates): Voice samples, annotated transcripts, or audio-related work
  • Professional experience in voice, linguistics, speech data, or speech evaluation and research

What the JD emphasized

  • native proficiency in Catalan
  • proficiency in English
  • strong auditory perception
  • demonstrated ability to handle multilingual audio content
  • demonstrated ability to transcribe audio with high accuracy
  • demonstrated ability to translate and localize text accurately
  • comfort providing high-quality voice recordings
  • strong comprehension skills
  • strong communication, interpersonal, analytical, detail-oriented, and organizational skills
  • commitment to developing AI that masters sophisticated multilingual audio capabilities

Other signals

  • training and refining Grok
  • curating and annotating high-quality audio data
  • improve AI's ability to handle speech modulation, accent variation, noise in real-world recordings, and multilingual audio processing