Harness and Platform Engineer, AI Safety and Security Engineering

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +5 · Remote

NVIDIA is seeking a Harness and Platform Engineer for their AI Safety and Security Engineering team. The role involves building and maintaining an agent harness and its underlying platform to support AI-powered tooling for vulnerability research and evaluation. Key responsibilities include developing the agent harness, building evaluation infrastructure for reproducible experiments, and ensuring environment and tooling consistency. The role requires strong Python engineering, familiarity with containerized environments, CI/CD, and agent frameworks, with a focus on reproducibility and reducing friction for researchers.

What you'd actually do

  1. Harness development: Design, implement, and maintain the agent harness.
  2. Evaluation infrastructure: Build the systems we use to run and reproduce experiments.
  3. Reproducibility: Own environments, tooling, and repeatable runs.
  4. Collaboration: Partner with security and evaluation engineers on their workflows.

Skills

Required

  • Python engineering
  • containerized environments
  • CI/CD
  • agent frameworks
  • LLM orchestration
  • ML infrastructure
  • evaluation harnesses

Nice to have

  • NVIDIA NeMo
  • security tooling
  • vulnerability research
  • internal tools development

What the JD emphasized

  • agent harness
  • platform
  • reproducible
  • repeatable
  • evaluation
  • tooling
  • environments

Other signals

  • AI Safety & Security Engineering team builds and evaluates AI-powered tooling
  • agent harness and the platform it runs on are the substrate everything else builds on
  • build that foundation
  • making the program's work possible, repeatable, and trustworthy
  • researcher runs an experiment, your platform decides whether the result can be reproduced tomorrow
  • NVIDIA NeMo agent tooling
  • keep the whole loop visible, from environment to run to result
  • environments that behave, tooling that stays versioned, and runs that repeat
  • debug it with the people who felt it first
  • treat researcher time as precious and remove friction wherever you find it