Senior Software Engineer, Aiops and Observability

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA

Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA, focusing on building AI agents and agentic workflows for monitoring, diagnosing, and optimizing products and services. This role involves implementing ML models for anomaly detection and root cause analysis, and developing AI-native observability tools.

What you'd actually do

  1. Lead the design, development, and deployment of AIOps & Observability platforms, including metrics, logs, traces, events, alerts, dashboards, and visualizations.
  2. Work with Data scientists to implement machine learning models for anomaly detection, forecasting, and root cause analysis on logs, metrics, and events. Handle large volumes of data and ensure data quality, security, and compliance.
  3. Develop AI agents and AI-native observability tools that help engineers detect, understand, and resolve production issues faster. Build agentic workflows that reason across logs, metrics, traces, events, alerts, topology, and incident history to support anomaly detection, forecasting, root cause analysis, automated debugging, and remediation recommendations.

Skills

Required

  • Observability tools (Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry)
  • AIOps tools (BigPanda, PagerDuty, Datadog)
  • Kubernetes, Nomad, Docker, microservices architectures
  • Streaming services (NATS, Kafka)
  • Programming languages (Go, Python, Java, C#)
  • Developing and operating observability platforms and solutions
  • Scalable data pipelines and instrumentation

Nice to have

  • Deep understanding of implementing Observability solutions to large scale on-prem Infrastructure and Networking.
  • Managing large scale Observability Platforms with LLMs & ML Models and building custom services to ingest billions of metrics and logs.
  • Developed unified cloud observability platform.
  • Using machine learning and Generative AI for predictive monitoring, incident diagnosis, summarization and correlation.
  • Proficiency in AI/ML systems, generative AI, or agentic AI frameworks.

What the JD emphasized

  • 5+ years of experience in developing and operating observability platforms and solutions
  • Experience with Kubernetes, Nomad, Docker, and microservices architectures
  • Experience with running large Observability platforms on BareMetal Infrastructure
  • Hands-on experience with managing large scale Observability Platforms with LLMs & ML Models
  • Demonstrated experience and expertise in using machine learning and Generative AI to develop solutions such as predictive monitoring, incident diagnosis, summarization and correlation.
  • Demonstrate proficiency in AI/ML systems, generative AI, or agentic AI frameworks.

Other signals

  • Develop AI agents and AI-native observability tools
  • Build agentic workflows that reason across logs, metrics, traces, events, alerts, topology, and incident history
  • implement machine learning models for anomaly detection, forecasting, and root cause analysis