Associate Director, AI Engineering - Live Operations

Verizon Verizon · Telecom · Basking Ridge, NJ +2

Associate Director of AI/ML Engineering focused on the production lifecycle, health, and optimization of deployed models in a large telecommunications company. The role involves leading a team to ensure AI applications are accurate, safe, and operationally efficient, with a strong emphasis on continuous operations, monitoring, evaluation, and automated deployment pipelines.

What you'd actually do

  1. Leading a talented, diverse team of AI scientists focusing on the continuous operations, health, and evolution of live production models.
  2. Monitoring and observing real-world model performance to proactively detect model drift, track latency, and manage computational expenses, including token costs for large language models.
  3. Designing and implementing continuous evaluation loops to systematically capture user feedback and ground-truth outcomes to assess model accuracy on a day-to-day basis.
  4. Managing incident response protocols, rapid hotfixing, and automated safety guardrails to block inappropriate inputs or outputs before they reach the customer.
  5. Establishing fallback architectures, such as routing users to human agents or traditional rule-based systems, to ensure a seamless experience if a model fails.

Skills

Required

  • Bachelor’s degree or four or more years of work experience.
  • Eight or more years of relevant experience required, demonstrated through one or a combination of work and/or military experience, or specialized training.

Nice to have

  • Master’s degree in statistics, mathematics, analytics, engineering, computer science, or another technical discipline.
  • Direct experience managing teams of data scientists or research-oriented talent in a production, operations, or LiveOps capacity.
  • In-depth working experience with deep learning frameworks (such as PyTorch), Generative AI (such as Large Language Models, Multimodal LLMs, and Alignment), and distributed computing systems.
  • Experience building and maintaining automated model monitoring, evaluation, and CI/CD pipelines in a cloud or edge production environment.
  • High-level understanding of, and the ability to explain, both the software engineering code and the underlying mathematical algorithms.
  • Proficiency in developing real-time safety filters, system guardrails, and automated fallback systems for conversational or predictive AI.
  • Coding proficiency in Python, Java, or R, alongside experience with AWS, Hadoop, or other distributed compute platforms.

What the JD emphasized

  • continuous operations
  • health
  • evolution of live production models
  • Monitoring and observing real-world model performance
  • model drift
  • latency
  • computational expenses
  • token costs
  • continuous evaluation loops
  • user feedback
  • ground-truth outcomes
  • model accuracy
  • incident response protocols
  • rapid hotfixing
  • automated safety guardrails
  • block inappropriate inputs or outputs
  • fallback architectures
  • routing users to human agents
  • traditional rule-based systems
  • seamless experience
  • Automating continuous deployment (CI/CD) pipelines
  • retrain models
  • test them against strict safety benchmarks
  • deploy updates without user disruption
  • technical thought leader
  • trusted advisor
  • broader data science teams
  • external vendors
  • senior leadership
  • Building an inclusive, highly engaged team culture
  • autonomy
  • technical excellence
  • rapid delivery of business outcomes
  • AI model interacts with live users
  • real work begins
  • deep data science
  • translate technical complexity into clear business value
  • inspiring leader
  • fast-paced environment
  • brings out the best in your team
  • automated model monitoring
  • evaluation
  • CI/CD pipelines
  • cloud or edge production environment
  • real-time safety filters
  • system guardrails
  • automated fallback systems
  • conversational or predictive AI

Other signals

  • industrialize our data science and AI capabilities
  • scale and maintain production AI
  • transition from enabling billions of automated, real-time predictions to trillions