AI Engineer, Recursive Self-improvement for Compute

AMD AMD · Semiconductors · Santa Clara, CA · Engineering

AI Engineer focused on building recursive self-improvement systems for compute workloads and hardware engineering. This role involves creating agentic loops for optimization, validation, and iteration, leveraging AI for program optimization, code generation, and performance engineering. It requires collaboration across AI research, hardware, and software teams.

What you'd actually do

  1. Build agentic and learning-driven optimization loops for compute workloads and hardware engineering workflows.
  2. Develop systems that generate, compile, test, benchmark, profile, and iterate on candidate improvements with minimal human intervention.
  3. Convert high-value engineering workflows into verifiable tasks with clear graders, harnesses, metrics, and failure feedback.
  4. Collaborate with AI researchers on reward design, reward shaping, reward hacking analysis, long-horizon optimization, and model improvement loops.
  5. Design feedback systems that accumulate useful data from successful attempts, failed attempts, profiler traces, benchmark results, and validation logs.

Skills

Required

  • Python
  • C++
  • C
  • HIP
  • CUDA
  • AI systems
  • performance engineering
  • hardware-aware optimization
  • agentic software development
  • optimization loops
  • validation
  • benchmarking
  • profiling
  • Reinforcement learning
  • post-training
  • reward modeling
  • evaluation methods
  • experiment platforms
  • data pipelines

Nice to have

  • ROCm
  • Triton
  • PyTorch
  • JAX
  • TensorFlow
  • distributed training/inference systems
  • automated program optimization
  • agentic coding systems
  • compiler optimization
  • math libraries
  • hardware design
  • simulation
  • formal verification
  • performance/power/area analysis
  • hardware/software co-design
  • production-quality evaluation platforms
  • experiment tracking
  • dashboards
  • leaderboards
  • Publications
  • open-source contributions
  • shipped systems in AI systems, GPU computing, compilers, RL, or hardware/software co-design

What the JD emphasized

  • correctness is non-negotiable
  • measurable outcomes: correctness, speed, efficiency, quality, reproducibility, and engineering leverage
  • minimal human intervention
  • clear graders, harnesses, metrics, and failure feedback
  • accumulate useful data from successful attempts, failed attempts, profiler traces, benchmark results, and validation logs
  • staged validation, caching, parallel execution, proxy metrics, and faster feedback paths
  • reproducible comparison and continuous improvement
  • evaluated with objective metrics
  • reliable experiment loops, benchmark harnesses, validation workflows, and correctness/performance evaluation pipelines
  • performance-oriented iteration

Other signals

  • recursive self-improvement systems
  • agentic software development
  • AI proposes improvements, verifies correctness, measures impact, learns from failures, and improves the next generation of compute workloads and platforms
  • AI-assisted program optimization, code generation, debugging, and automated repair
  • Agentic workflows that use tools, tests, simulators, profilers, and structured feedback to improve over time