Data Annotation Specialist, Generalist

Cohere Cohere · AI Frontier · Canada · Data Quality (Contract)

Cohere is seeking a Data Annotation Specialist to evaluate, rank, and correct model outputs, stress-test models for failure modes, create datasets, build rubrics, and annotate multimodal data. This role is crucial for improving Large Language Models' performance and ensuring data quality.

What you'd actually do

  1. Evaluate and rank model outputs
  2. Stress-test and break models
  3. Create datasets
  4. Build and apply rubrics and taxonomies
  5. Annotate and correct multimodal data

Skills

Required

  • AI data annotation
  • LLM evaluation
  • content moderation
  • research
  • quality assurance
  • preference ranking
  • applying detailed guidelines
  • contextual and sociocultural judgment
  • sensitivity to nuance, tone, and register
  • reasoning in ambiguous cases
  • identifying model failure modes
  • probing models adversarially
  • written English proficiency
  • reading comprehension
  • justifying evaluations
  • writing clear prompts and exemplars
  • attention to detail
  • commitment to accuracy
  • consistency across high-volume tasks
  • working with annotation platforms
  • structured data formats (JSON, CSV/TSV, Markdown, XML, YAML)
  • remote work execution
  • time management
  • using new tools
  • independent work

Nice to have

  • fluency in another language

What the JD emphasized

  • broad backgrounds that span multiple consumer-facing or personal domains
  • judgment-driven role, not passive data entry
  • evaluating, stress-testing, and improving our models on English-language data spanning multiple modalities (text, image, and structured formats such as JSON, CSV/TSV, and Markdown)
  • strong analytical skills
  • assess which responses best conform to project guidelines for accuracy, helpfulness, tone, and safety, and writing clear justifications for your judgments
  • Probe models adversarially to surface failure modes, unsafe behavior, and capability gaps, and document reproducible cases that engineering and research teams can act on
  • Author high-quality prompts, responses, and exemplars to build training and evaluation datasets, following detailed specifications and editing machine-written or human-written outputs to standard
  • Contribute to the design of grading criteria and rubrics, then apply them consistently to produce structured, high-quality annotations across task types
  • Label, audit, and rectify inaccuracies across text, image, and structured data, maintaining a high standard of data integrity and accuracy
  • Calibrate and maintain consistency
  • Adapt to experimental work
  • Report on model performance
  • 1+ years of experience in AI data annotation, LLM evaluation, content moderation, research, or a related analytical role, with exposure to quality assurance, and/or preference ranking
  • Experience applying detailed guidelines to complex and often ambiguous content, with strong contextual and sociocultural judgment, sensitivity to nuance, tone, and register, and the ability to reason well in cases where there is no single correct answer
  • Comfort with ambiguity: a willingness to flag unclear or uncovered edge cases rather than guess, and to work productively on novel, experimental tasks whose definitions are still evolving
  • A sharp, curious eye for inconsistencies, subtle errors, and model failure modes, including the instinct to probe models adversarially and surface where they break
  • Familiarity with how large language models behave, such as hallucination, sycophancy, and instruction-following gaps, is a plus
  • Excellent command of written English and strong reading comprehension, with the ability to clearly justify your evaluations, including why an output is correct or incorrect, high-quality or low-quality, and to write clear prompts and exemplars that explain your reasoning
  • Strong attention to detail and commitment to accuracy, with the ability to maintain consistency across high-volume and monotonous tasks
  • Comfort working with annotation platforms and structured formats such as JSON, CSV/TSV, Markdown, XML, and YAML
  • Strong execution in a remote environment, including good time management, comfort using new tools, and the ability to work independently in a global, asynchronous team

Other signals

  • data annotation
  • model evaluation
  • dataset creation
  • multimodal data