Senior AI Engineer - Live Operations

Verizon Verizon · Telecom · Basking Ridge, NJ +2

Senior AI/ML Engineer focused on the Live Operations of deployed AI models, ensuring stability, safety, and efficiency. The role involves maintaining real-time monitoring, detecting model degradation, implementing continuous evaluation, deploying safety filters and guardrails, integrating fallback systems, developing CI/CD pipelines for model updates, resolving production incidents, and collaborating on scaling backend architectures. This position is critical for industrializing AI capabilities and scaling production AI systems in a telecommunications environment.

What you'd actually do

  1. Building and maintaining real-time monitoring and observability pipelines to track system throughput, latency, and token consumption costs for large language models.
  2. Monitoring production model behavior to identify data drift and model degradation, ensuring our systems maintain high predictive accuracy over time.
  3. Implementing automated continuous evaluation loops that capture user feedback and ground-truth telemetry to dynamically assess performance.
  4. Deploying real-time safety filters, system guardrails, and input/output moderators to mitigate model hallucinations and prevent inappropriate content generation.
  5. Resolving production model incidents as a senior escalation contact, diagnosing issues with live models, and shipping immediate hotfixes or prompt adjustments.

Skills

Required

  • advanced technical expertise in data science, system orchestration, and software engineering
  • maintain the stability, safety, and efficiency of our deployed artificial intelligence models
  • health of machine learning systems
  • real-time monitoring and observability pipelines
  • production model behavior monitoring
  • automated continuous evaluation loops
  • real-time safety filters, system guardrails, and input/output moderators
  • fallback systems
  • automated continuous integration and continuous deployment pipelines
  • production model incidents resolution
  • diagnosing issues with live models
  • shipping immediate hotfixes or prompt adjustments
  • scaling backend architectures
  • technical alignment
  • secure, ethical AI standards
  • Python, Java, or C++
  • data structures
  • cloud systems such as AWS, Google Cloud, or Azure
  • distributed compute architectures
  • database design
  • real-time streaming technologies such as Kafka or Spark

Nice to have

  • Master's degree in computer science, data science, electrical engineering, mathematics, or another highly technical discipline
  • deploying, monitoring, and debugging complex machine learning or deep learning models in large-scale production environments
  • MLOps frameworks and automated pipeline tools, including Airflow, Kubeflow, or MLflow
  • Deep experience developing safety guardrail middleware, system monitoring dashboards, and telemetry systems for Generative AI applications

What the JD emphasized

  • maintain the stability, safety, and efficiency of our deployed artificial intelligence models
  • health of machine learning systems once they are exposed to real-world datasets and production traffic
  • industrialize our data science and AI capabilities
  • build and scale production AI systems
  • trillions of real-time inferences
  • real-time monitoring and observability
  • production model behavior
  • automated continuous evaluation loops
  • real-time safety filters, system guardrails, and input/output moderators
  • mitigate model hallucinations and prevent inappropriate content generation
  • fallback systems
  • automated continuous integration and continuous deployment pipelines
  • Resolving production model incidents
  • diagnosing issues with live models
  • shipping immediate hotfixes or prompt adjustments
  • scale backend architectures
  • foster technical alignment
  • champion secure, ethical AI standards
  • deploying, monitoring, and debugging complex machine learning or deep learning models in large-scale production environments
  • MLOps frameworks and automated pipeline tools
  • Deep experience developing safety guardrail middleware, system monitoring dashboards, and telemetry systems for Generative AI applications

Other signals

  • industrialize our data science and AI capabilities
  • build and scale production AI systems
  • transition from millions of automated predictions to trillions of real-time inferences