Senior Devops Engineer, Cloud Simulation Infrastructure

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA

Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets, enabling high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF). The role involves deploying multi-GPU pipelines for structural validation, AI-driven runtime behavioral testing, and automated asset remediation, as well as owning cloud infrastructure, artifact pipelines, observability, and CI/CD.

What you'd actually do

  1. Deployment: Deploy full Isaac Sim runtimes within GPU-aware NVCF containers. Manage container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.
  2. Deploy Runtime Validation: Architect scalable execution layers to conduct runtime behavior-based testing (e.g., drop/grasp tests). Deploy rule-based systems or AI based systems for automated pass/fail grading.
  3. Deploy Automated Remediation: Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results to ensure quality standards.
  4. Cloud Infrastructure Ownership: Scale execution from single-workstation validation to massive, multi-GPU cloud environments. Optimize for performance, addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.
  5. Observability & CI/CD: Establish robust CI/CD, cluster verification, and monitoring pipelines. Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.

Skills

Required

  • 8+ years of professional experience working on DevOps and/or cloud simulation
  • Extensive experience in production-grade DevOps, SRE, or Infrastructure Engineering, with a focus on GPU-backed cloud services
  • Proven expertise in container orchestration (Kubernetes/Docker) and CI/CD pipeline development
  • Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation
  • Proficiency in Python and systems scripting for test orchestration and pipeline automation
  • Strong ability to design and maintain distributed job lifecycle services (submit/poll/fetch/cancel) and handle asynchronous failure states
  • Ability to diagnose and solve distributed network bottlenecks, including gRPC and function-to-function communication

Nice to have

  • Direct experience deploying services on NVCF (NVIDIA Cloud Functions) or DGX Cloud
  • Deep familiarity with Isaac Sim, Omniverse, USD, or Sensor RTX workflows
  • Background in robotics simulation, physical AI, or large-scale content creation pipelines
  • Experience building "self-healing" or automated remediation workflows
  • Experience with cluster verification frameworks, stress testing, and deployment validation at scale

What the JD emphasized

  • AI-driven runtime behavioral testing
  • AI based systems for automated pass/fail grading
  • AI-based pipeline that intercepts failures
  • automated asset fixes
  • automated remediation
  • AI/ML inference
  • automated testing frameworks

Other signals

  • Deploy AI-based systems for automated pass/fail grading
  • Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results
  • Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation