Security Engineer, Detection and Response

Writer Writer · AI Frontier · San Francisco, CA · Engineering, product & design

Security engineer focused on detection and response for AI infrastructure, including training data, model deployments, and inference endpoints. The role involves building automated response capabilities, leading incident response coordination, and proactive threat hunting within GPU clusters and distributed training environments. It requires collaboration with AI Security research, Cloud Infrastructure, and Software Security Engineering teams to protect enterprise-grade LLMs and AI platforms.

What you'd actually do

  1. Design and implement detection strategies that identify AI-specific threats including prompt injection, model extraction, data poisoning, adversarial examples, and unauthorized access to training datasets or model weights across our distributed infrastructure
  2. Build automated response playbooks and orchestration workflows that contain threats without human intervention, creating self-healing security systems that reduce mean time to response from hours to minutes while automatically remediating compromised inference endpoints
  3. Lead security incident response coordination across all teams (Cloud, AppSec, Enterprise, AI Security) when AI infrastructure or models are compromised, conducting forensic investigations on training pipeline attacks and model manipulation attempts while drafting clear incident communications for engineering and executive leadership
  4. Hunt proactively for sophisticated threats across GPU clusters and training infrastructure by analyzing model outputs for signs of compromise, reproducing AI-specific vulnerabilities from security research, and identifying visibility gaps in distributed training environments before adversaries exploit them
  5. Build detection-as-code frameworks with version control and automated deployment, onboard telemetry from AI training infrastructure and inference endpoints, and create dashboards that track model security metrics, GPU utilization patterns, and access to sensitive research data

Skills

Required

  • 3-5+ years in security operations, detection engineering, or incident response
  • Proven track record of identifying and stopping sophisticated attacks in production environments
  • Securing AI/ML infrastructure, high-performance computing environments, or other distributed systems at scale
  • Strong programming skills in Python, KQL, SPL, or similar languages
  • Experience with SIEM platforms, detection technologies, and forensic investigation techniques
  • Demonstrated ability to build detection for novel attack techniques
  • Conduct forensics in complex distributed environments
  • Self-directed execution mindset
  • Track record of securing high-value intellectual property
  • Automating incident response in complex environments
  • Identifying critical security gaps through proactive threat hunting

Nice to have

  • Collaborate cross-functionally as the operational security partner for all teams
  • Translating AI Security's threat research into production detections
  • Monitoring Cloud Infrastructure's GPU clusters for threats
  • Detecting customer-impacting incidents for Software Security Engineering
  • Enabling responsible AI development through security guardrails

What the JD emphasized

  • AI-specific threats
  • automated response
  • incident response coordination
  • training datasets or model weights
  • AI infrastructure
  • distributed training environments
  • AI Security research
  • AI platform
  • securing systems that are fundamentally different
  • novel threats that don't exist in textbooks yet
  • AI security engineering at scale
  • securing AI/ML infrastructure
  • novel attack techniques

Other signals

  • Defending AI infrastructure
  • Building detection systems for AI-specific threats
  • Automated response for AI systems
  • Securing training data and model deployments
  • Incident response for AI infrastructure and models