ML Runtime Optimization Engineer

Applied Intuition Applied Intuition · Robotics · Sunnyvale, CA · SDS Software Engineering

ML Runtime Optimization Engineer focused on optimizing ML model performance and inference on embedded runtime environments for physical AI applications in robotics and autonomous systems.

What you'd actually do

  1. Drive ML performance optimization on multiple technologies for on-road and off-road ADAS / AD stacks targeting deployment on a variety of embedded compute platforms
  2. Develop compute usage strategies to optimize efficiency and latency of model inference for compute boards selected by our customers
  3. Work on model pruning and quantization, and support deployment on memory constrained platforms
  4. Collaborate closely with ML engineers and software developers on technical efforts to find and optimize efficient model architecture solutions
  5. Set up methodologies to profile the model performance on target embedded compute platforms and identify performance bottlenecks as part of stack integration

Skills

Required

  • ML accelerators
  • GPU
  • CPU
  • SoC architecture and micro-architecture
  • software development skills
  • embedded programming
  • profiling and optimizing model performance on embedded compute platforms
  • deep learning frameworks (e.g., PyTorch, JAX, ONNX, etc.)

Nice to have

  • M.Sc or PhD in a ML related area
  • Built an ML optimization framework from scratch
  • Deployed ML solutions to embedded chips for real time robotics applications

What the JD emphasized

  • embedded programming
  • profiling and optimizing model performance on embedded compute platforms

Other signals

  • ML performance optimization
  • inference efficiency and latency
  • model pruning and quantization
  • embedded compute platforms