Safeguards Enforcement Analyst, User Well-being

Anthropic Anthropic · AI Frontier · San Francisco, CA · Safeguards (Trust & Safety)

This role supports the design and deployment of mental health guardrails for AI systems, focusing on detection systems, review queues, and intervention evaluation. It involves translating clinical guidance and data analysis into actionable changes for AI response and monitoring. The role partners with Engineering and Data Science to build and tune detection models, monitors system performance, reviews flagged content, and supports the development of in-product features. It requires experience in trust & safety or policy, with a focus on mental health harms, and proficiency in data analysis tools. Experience with generative AI products and LLM-based classification systems is preferred.

What you'd actually do

  1. Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
  2. Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems, including threshold-setting and precision and recall tradeoffs
  3. Monitor how interventions and detection systems perform over time
  4. Review flagged content to drive enforcement and policy improvements
  5. Support the development of in-product features that connect users to crisis resources, working with Product, Legal, and external partners on referral pathways and user-facing content

Skills

Required

  • Experience in trust & safety, product policy, content moderation, or a related field, with direct exposure to mental health, suicide and self-harm, or related well-being harm areas
  • Experience designing or running experiments, evaluations, or measurement studies to determine whether an intervention worked
  • Experience translating policy definitions into measurable form — the rubrics, review guidelines, or classification criteria, whether applied by human reviewers or automated systems
  • Experience managing or coordinating content review operations, including quality assurance and workflow management
  • Proficiency in SQL and/or other data analysis tools to measure intervention efficacy, monitor workflow health, and surface policy gaps
  • Experience working with generative AI products, including writing effective prompts for content review, classification, or evaluation
  • Experience turning open questions and data into concise and insightful analysis
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders
  • Understanding of the challenges involved in implementing product policies at scale in the content moderation space
  • Sound judgment in ambiguous, high-consequence cases, and comfort making a call and escalating appropriately when the available signal is incomplete

Nice to have

  • Subject matter expertise in mental health, whether developed in academia, clinical practice, crisis intervention, trust & safety, or other related settings
  • Experience building or evaluating LLM-based classification systems
  • Experience using agentic tools (e.g. Claude Code) to scale analysis or automate recurring work
  • Experience working within crisis support

What the JD emphasized

  • mental health guardrails
  • detection models
  • intervention efficacy
  • content moderation

Other signals

  • design and deployment of mental health guardrails
  • iterating on detection systems
  • managing review queues
  • evaluating new interventions
  • monitoring existing ones
  • translating expert clinical guidance, internal data analyses, and legal or engineering constraints into concrete changes to how we detect, review, and respond