Senior/staff Software Engineer, AI Agent Infrastructure

Nuro Nuro · Robotics · CA · AI Platform

Senior/Staff Software Engineer to build and scale AI agent infrastructure at Nuro, focusing on closed-loop evaluation, agent platform capabilities (orchestration, sandboxing, tool use, observability), and automating the research loop for autonomous driving models. The role involves ensuring trustworthiness and autonomy of AI agents operating on production systems and infrastructure.

What you'd actually do

  1. Build the closed-loop measurement layer that tells us, per workflow, whether agent output is accepted, reverted, or overridden — and use it to decide where autonomy expands and where it gets pulled back.
  2. Take the autoresearch loop from assisted to unattended for a bounded class of experiments, including the eval and confidence machinery required to run it without a human in the loop.
  3. Design the isolation and permissioning model that lets agents act on production repositories and infrastructure with an auditable record of what they did and why.

Skills

Required

  • 5+ years of software engineering experience (or 4+ with a Master's)
  • Deep, current taste in LLM research
  • Understanding of model training (data, tokenization, architecture, pretraining dynamics)
  • Understanding of post-training stack (supervised fine-tuning, preference optimization, RL)
  • Reasoning about model behavior based on training decisions

Nice to have

  • Experience with agent orchestration
  • Experience with evaluation frameworks
  • Experience with model serving infrastructure
  • Experience with research automation

What the JD emphasized

  • closed-loop evaluation system
  • rigorous closed-loop evaluation system
  • trustworthy enough to run unattended
  • most rigorous closed-loop evaluation system for AI work anywhere
  • trust follows from measurement
  • agent platform
  • orchestration, sandboxing and isolation, tool and skill frameworks, memory, identity and permissioning, and the gateway and observability layer underneath
  • containment and auditability
  • autoresearch infrastructure
  • automation of the research loop itself
  • trusting the measurement
  • surviving experiments that take days
  • spending finite research compute wisely
  • producing proposals a skeptical researcher can audit and reject
  • Deep, current taste in LLM research
  • understand how a model is trained from scratch — data, tokenization, architecture, pretraining dynamics, the full post-training stack of supervised fine-tuning, preference optimization, and RL — and you can reason about what a training decision does to model behavior

Other signals

  • building agent infrastructure
  • closed-loop evaluation systems
  • autonomous agents
  • research automation