Senior Software Engineer - Local AI

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +1

Senior Software Engineer focused on optimizing and deploying AI models for client-side execution on resource-constrained devices (RTX and DGX PCs). This role involves improving performance through techniques like quantization, distillation, and pruning, and collaborating with various teams to foster the local AI ecosystem.

What you'd actually do

  1. Partnering with NVIDIA software, research, architecture, and product teams to align strategies and technical needs for fostering the ecosystem of AI on RTX and DGX PCs.
  2. Collaborate closely with industry partners to advance AI across critical domains—including graphics, web browsers, and edge devices—by driving innovation in both open and closed source technologies with emphasis on system level support.
  3. Improving performance on current and next-generation GPU architectures by conducting in-depth analysis and end-to-end optimization of AI models, data processing pipelines, and inference runtime features.
  4. Identifying, evaluating, and implementing compute and memory optimization techniques—such as quantization, distillation, and pruning—for large AI models; fine-tuning and compressing models to fit edge devices.

Skills

Required

  • C++
  • data structures
  • algorithms
  • AI inferencing pipelines
  • ML/DL frameworks
  • ONNX RT
  • PyTorch
  • Tensor RT
  • llama.cpp
  • vLLM

Nice to have

  • Machine Learning
  • Deep Neural Networks
  • Generative AI
  • open-source projects
  • end-to-end products
  • lower-level system/GPU programming
  • CUDA
  • high-performance systems
  • ONNX RT
  • DirectX
  • PyTorch
  • TensorRT
  • Vulkan
  • llama.cpp

What the JD emphasized

  • 5+ years of experience
  • Excellent C++ programming and debugging skills
  • proficiency in AI inferencing pipelines and applications using ML/DL frameworks
  • Strong analytical and problem-solving abilities
  • Outstanding written and oral communication skills

Other signals

  • running AI models locally
  • client-side AI
  • inference runtime features
  • optimization techniques