Cyber Operations Strategist, Critical Harm Operations

OpenAI OpenAI · AI Frontier · San Francisco, CA · User Operations

OpenAI is seeking a senior cybersecurity practitioner and operations strategist to enhance the quality, scalability, and technical rigor of its Cyber Operations. The role involves resolving complex dual-use questions, evolving standard operating procedures (SOPs), improving reviewer capabilities, and developing practical tools and automations. The individual will drive the operating model, serve as a senior cyber expert for critical decisions across various OpenAI products, and translate policy into clear decision rules, training, and tooling requirements. Key responsibilities include building durable operating systems and quality loops, raising the capability of human reviewers and vendors, diagnosing root causes of operational issues, and building solutions like SQL analyses, scripts, dashboards, LLM eval workflows, and lightweight automations. The role requires extensive hands-on cybersecurity experience, deep understanding of attacker tradecraft, and the ability to translate complex judgment into actionable guidance and operational improvements. Experience with trust and safety, platform abuse, cyber misuse of AI, or LLM safety is also valuable.

What you'd actually do

  1. Drive the Cyber Operations operating model across domain priorities, SOPs, escalation paths, quality health, vendor capability, roadmap inputs and help inform trusted access strategies.
  2. Serve as the senior cyber expert for complex or high-risk decisions across ChatGPT, API, Codex, agents, and emerging product surfaces.
  3. Translate policy ambiguity, quality misses, appeals, and reviewer disagreement into clear decision rules, calibration examples, training, and tooling requirements.
  4. Build durable operating systems and quality loops: golden sets, holdouts, double-labeling, adjudication, error taxonomies, reviewer calibration, and automation evaluations.
  5. Build hands-on solutions—SQL analyses, scripts, dashboards, LLM eval workflows, evidence enrichment, routing logic, and lightweight automations—that improve decision quality and reduce manual effort.

Skills

Required

  • 8+ years of hands-on cybersecurity experience in offensive security, threat intelligence, incident response, security research, red teaming, application security, DFIR, malware analysis, or a related field.
  • Ability to reason deeply about attacker tradecraft, vulnerability exploitation, credential abuse, malware, persistence, evasion, exfiltration, cloud or identity abuse, and ambiguous dual-use activity.
  • Experience building or improving high-stakes operations, reviewer programs, QA systems, escalation workflows, or vendor/BPO programs.
  • Ability to translate complex cyber and policy judgment into reviewer-usable SOPs, decision trees, training, and concise written recommendations.
  • Comfort using SQL, Python, C/C++, JavaScript, PowerShell, and Bash, APIs, LLM tooling, or automation to solve operational problems.
  • Understanding of human-in-the-loop automation, evaluations, monitoring, holdouts, and fallback paths for sensitive workflows.
  • Ability to operate independently in ambiguity, communicate clearly across technical and non-technical audiences, and bring sound judgment and discretion when handling sensitive material.
  • Experience with trust and safety, platform abuse, cyber misuse of AI systems, or LLM safety.

Nice to have

  • Experience with golden sets, classifier or prompt evaluations, and reviewer-quality programs.
  • Experience enabling global vendor reviewer operations.

What the JD emphasized

  • senior cybersecurity practitioner
  • hands-on cyber judgment
  • systems-level operating design
  • resolve the hardest dual-use questions
  • durable improvement in the operating model
  • senior cyber expert
  • policy ambiguity
  • quality misses
  • reviewer disagreement
  • durable operating systems
  • automation evaluations
  • diagnose root causes
  • prioritize high-leverage fixes
  • improve decision quality
  • reduce manual effort
  • hands-on cybersecurity experience
  • reason deeply about attacker tradecraft
  • built or improved high-stakes operations
  • turn complex cyber and policy judgment
  • reviewer-usable SOPs
  • human-in-the-loop automation
  • evaluations
  • monitoring
  • holdouts
  • fallback paths
  • operate independently in ambiguity
  • sound judgment and discretion
  • handling sensitive material
  • cyber misuse of AI systems
  • LLM safety
  • golden sets
  • classifier or prompt evaluations
  • reviewer-quality programs

Other signals

  • building enforcement systems
  • operationalizing changes
  • improving decision quality
  • reducing manual effort
  • LLM eval workflows