Evaluation and ML Systems Engineer, AI Safety and Security Engineering

NVIDIA NVIDIA · Semiconductors · Santa Clara, CA +5 · Remote

This role focuses on building and maintaining the evaluation infrastructure for AI safety and security tooling. The primary responsibility is to design metrics, establish protocols, and ensure the reproducibility and traceability of AI model evaluations, playing a critical role in measuring program progress and validating findings.

What you'd actually do

  1. Evaluation infrastructure: Build the benchmarking and reproducibility systems we depend on.
  2. Metrics and protocols: Define the metrics and protocols we measure against.
  3. Traceability: Map every result to the code and runs that produced it.
  4. Evidence discipline: Keep findings reviewable and conclusions traceable.

Skills

Required

  • Python engineering
  • experiment tracking
  • data pipelines
  • ML engineering
  • evaluation experience

Nice to have

  • security evaluation
  • evaluating security tooling or pipelines
  • agentic systems
  • measuring agent or LLM behavior
  • Contributions to public benchmarks or evaluation frameworks

What the JD emphasized

  • Evaluation experience: Designing benchmarks, metrics, and statistically sound comparisons for ML systems.
  • Measurement rigor: A careful, skeptical approach to metrics, baselines, and claims.
  • Security evaluation: Exposure to evaluating security tooling or pipelines.
  • Agentic systems: Experience measuring agent or LLM behavior.

Other signals

  • measurement is not a support function
  • design the metrics, choose the baselines, and write the protocols
  • build the infrastructure that keeps every result tied to the run that produced it
  • automate the boring parts of measurement