ML Systems Research Engineer, RL / Inference / Agent Systems

AMD AMD · Semiconductors · Santa Clara, CA · Engineering

This role focuses on building the reinforcement learning, inference, and evaluation infrastructure for AI-for-engineering systems. The goal is to enable agents and models to improve real engineering workflows by managing attempts, evaluating correctness, measuring performance, handling long-latency rewards, and feeding results back into model and agent improvement. The role involves working across compute optimization, hardware engineering automation, verification, and simulation, with an emphasis on scalable ML systems for practical, repeatable, and useful production engineering.

What you'd actually do

  1. Build RL and inference systems for agentic engineering workflows, including job orchestration, sampling, scoring, caching, experiment tracking, and reproducible evaluation.
  2. Develop infrastructure for long-horizon and high-latency reward tasks where validation can take minutes to hours.
  3. Design staged rewards, proxy graders, sliced evaluation paths, retry strategies, and uncertainty-aware evaluation methods.
  4. Support optimization workflows with systems for candidate generation, benchmark execution, correctness checking, profiler feedback, reward modeling, and model-level improvement.
  5. Partner with AI research scientists on reward hacking research, reward shaping, metareasoning, and post-training methods for engineering tasks.

Skills

Required

  • Python
  • ML frameworks (PyTorch, JAX, TensorFlow)
  • ML systems
  • RL infrastructure
  • inference services
  • agent frameworks
  • evaluation platforms
  • distributed experimentation systems
  • model inference
  • batching
  • sampling
  • latency
  • throughput
  • observability
  • reliability tradeoffs
  • experiment design
  • evaluation pipelines
  • clear metrics
  • logs
  • reproducibility
  • statistical discipline

Nice to have

  • Reinforcement learning
  • RLHF
  • GRPO
  • preference optimization
  • reward modeling
  • reward shaping
  • post-training systems
  • LLM agents
  • tool-use systems
  • code generation
  • automated program repair
  • compiler optimization
  • benchmark-driven development
  • distributed systems
  • job orchestration
  • Kubernetes
  • Ray
  • Slurm
  • workflow engines
  • data pipelines
  • large-scale experiment management
  • GPU systems
  • ROCm/HIP
  • CUDA
  • profiling
  • kernel benchmarking
  • model serving
  • distributed training/inference
  • hardware engineering workflows
  • design
  • verification
  • firmware
  • simulation
  • performance analysis
  • Publications or shipped systems in ML systems, RL, inference optimization, AI infrastructure, or hardware/software co-design

What the JD emphasized

  • scalable ML systems
  • long-horizon and high-latency reward tasks
  • uncertainty-aware evaluation methods
  • scalable inference and tool-use pipelines
  • Standardize datasets, eval definitions, run logs, leaderboards, failure taxonomies, and data collection for future training.
  • Reinforcement learning and post-training infrastructure for tool-using agents.
  • Inference systems for LLMs and agents, including latency, throughput, batching, sampling, reliability, and observability.
  • Evaluation systems for tasks with expensive, delayed, mixed, or sparse rewards.
  • Distributed experimentation, job orchestration, caching, data pipelines, dashboards, and reproducible run management.
  • Integration with external tools such as compilers, profilers, simulators, validation systems, benchmark harnesses, and ticketing or knowledge systems.
  • Experience building ML systems, RL infrastructure, inference services, agent frameworks, evaluation platforms, or distributed experimentation systems.
  • Strong understanding of model inference, batching, sampling, latency, throughput, observability, and reliability tradeoffs.
  • Ability to design experiments and evaluation pipelines with clear metrics, logs, reproducibility, and statistical discipline.
  • Experience with reinforcement learning, RLHF, GRPO, preference optimization, reward modeling, reward shaping, or post-training systems.
  • Experience with LLM agents, tool-use systems, code generation, automated program repair, compiler optimization, or benchmark-driven development.
  • Experience with distributed systems, job orchestration, Kubernetes, Ray, Slurm, workflow engines, data pipelines, or large-scale experiment management.
  • Familiarity with GPU systems, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.
  • Publications or shipped systems in ML systems, RL, inference optimization, AI infrastructure, or hardware/software co-design are valued.

Other signals

  • Reinforcement learning
  • Inference systems
  • Agent systems
  • Evaluation infrastructure
  • Scalable ML systems