Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA

Senior Perception Engineer for NVIDIA's autonomous driving division, focusing on designing, implementing, and productizing 3D obstacle perception systems using advanced deep learning techniques (CNNs, transformers, multi-modal, vision-language). The role involves developing production-grade models, defining data strategies, collaborating with cross-functional teams, and optimizing for safety, latency, and robustness.

What you'd actually do

  1. Develop and improve the technical design, architecture, and roadmap for 3D obstacle perception to support end-to-end autonomous driving functionalities, leveraging state-of-the-art CNN and transformer-based architectures where appropriate.
  2. Design and implement advanced 3D perception models using multi-camera inputs and/or multi-sensor fusion (camera, radar, lidar) for obstacle detection and tracking, including opportunities to explore BEV and transformer-based 3D perception.
  3. Build efficient, production-grade deep learning models: define objectives with the team, select and prototype architectures, run experiments, and follow best practices for training and evaluation, using techniques such as large-scale pretraining, distillation, and parameter-efficient fine-tuning (e.g., LoRA).
  4. Help define and maintain KPI frameworks to quantify perception performance; analyze large-scale real and synthetic datasets to identify failure modes and systematically improve accuracy, robustness, and efficiency, incorporating approaches like self-supervised and representation learning when beneficial.
  5. Contribute to the data strategy for perception: specify data and labeling requirements, help prioritize data collection and annotation, and collaborate with data and ground-truth teams, including model-assisted workflows (e.g., active learning, auto-labeling, vision-language models (VLMs)) and model-in-the-loop tooling.

Skills

Required

  • PhD with 4+ years, MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field.
  • Hands-on experience developing deep learning–based perception or closely related systems for complex real-world problems, with strong proficiency in frameworks such as PyTorch and a track record of taking models from prototype to production.
  • Proven experience in data-driven development, including close collaboration with data, labeling, and ground-truth teams on data strategy, labeling quality, and iterative model improvement.
  • Strong programming skills in Python and/or C++, with experience building reliable, high-performance, production-quality software.
  • Excellent communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.

Nice to have

  • Experience designing and deploying perception solutions for autonomous driving or robotics using camera-based deep learning at scale.
  • Hands-on experience architecting and deploying DNN-based perception pipelines on embedded or real-time platforms, including optimization for latency, memory, and compute constraints, and experience with modern architectures such as CNNs and transformers, plus familiarity with techniques like large-scale pretraining, parameter-efficient fine-tuning (e.g., LoRA), or vision-language models (VLMs).
  • Strong publication record or recognized contributions in deep learning, computer vision, or autonomous systems at leading conferences/journals (e.g., CVPR, ICCV, NeurIPS, IROS).
  • Deep understanding of 3D computer vision fundamentals, including camera modeling and calibration (intrinsic and extrinsic), multi-view geometry, and 3D representations, ideally with experience applying these concepts in transformer-based 3D or BEV perception pipelines.
  • Experience with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated components.

What the JD emphasized

  • production-grade deep learning models
  • production-quality software
  • data strategy
  • labeling requirements
  • perception performance
  • obstacle detection and tracking
  • autonomous driving

Other signals

  • production-grade deep learning models
  • state-of-the-art CNN and transformer-based architectures
  • multi-sensor fusion
  • large-scale pretraining
  • parameter-efficient fine-tuning
  • data strategy for perception
  • model-assisted workflows
  • production-quality software