Engineering Manager - Production Services Infrastructure (remote)

CrowdStrike CrowdStrike · Enterprise · CA · Remote

Engineering Manager to lead a team responsible for CrowdStrike's global production infrastructure, focusing on large-scale distributed systems and implementing AI/ML solutions for optimization and automation. The role involves technical strategy, team leadership, and driving efficiency through AI.

What you'd actually do

  1. Lead and scale an engineering team responsible for CrowdStrike's global production infrastructure
  2. Drive technical strategy for large-scale distributed systems supporting millions of security endpoints
  3. Implement AI/ML solutions for predictive maintenance, anomaly detection, and automated remediation
  4. Build and mentor a world-class team while establishing career development frameworks
  5. Drive automation initiatives reducing manual operational overhead by thousands of hours annually

Skills

Required

  • Leading engineering teams
  • Technical strategy for distributed systems
  • ML Engineering
  • AI applications for infrastructure optimization
  • Kubernetes
  • Virtualization
  • Cloud platforms
  • Bare metal orchestration
  • Infrastructure automation
  • CI/CD systems
  • Configuration management
  • Observability platforms
  • Log aggregation
  • Event-driven architecture
  • Presenting technical strategies to executive leadership
  • Project management

Nice to have

  • Cybersecurity infrastructure
  • Threat detection platforms
  • Security operations
  • Agentic AI
  • LLM applications
  • Advanced ML model deployment
  • Self-service platforms
  • Federated platform architectures
  • Statistical analysis
  • Advanced certifications in cloud platforms, ML, or infrastructure technologies
  • Platform improvements reducing time-to-production

What the JD emphasized

  • 10+ years of engineering experience with 5+ years leading technical teams of 5+ engineers
  • Proven track record delivering $100M+ in cost savings through infrastructure optimization
  • Deep expertise in large-scale distributed infrastructure (5M+ servers/endpoints preferred)
  • Strong background in ML Engineering and AI applications for infrastructure optimization
  • Expert-level experience with Kubernetes, virtualization, cloud platforms, and bare metal orchestration
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.

Other signals

  • AI/ML solutions for predictive maintenance, anomaly detection, and automated remediation
  • large scale distributed systems
  • AI-native platform