Embedded AI Engineer – Android Automotive (on-device Intelligence)

Applied Intuition Applied Intuition · Robotics · Sunnyvale, CA · Vehicle OS

This role focuses on deploying and optimizing on-device ML inference and learning systems for Android Automotive. It involves implementing multimodal LLMs, integrating models with specific runtimes, profiling for strict performance budgets, and designing safety guardrails for model outputs within embedded constraints.

What you'd actually do

  1. Deploy and run production-grade ML inference and learning systems on Android Automotive (AAOS)
  2. Implement on-device multimodal LLMs, including schema design and safe dispatch to local vehicle APIs
  3. Integrate models using TensorFlow Lite, ONNX Runtime, or specialized vendor SDKs
  4. Profile and optimize models for strict latency, memory, power, and thermal budgets
  5. Design safety boundaries and guardrails for model outputs, including tool-call allowlists and fallback logic

Skills

Required

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field
  • 3+ years of experience shipping ML inference on embedded, mobile, or automotive platforms
  • Strong proficiency in C++ and experience with native Android integration (JNI)
  • Expertise in model optimization techniques such as quantization, pruning, and compilation
  • Experience integrating LLM function calling or tool execution with structured outputs
  • Hands-on experience with Android system services or Android Automotive OS (AAOS)
  • Deep understanding of edge constraints including real-time behavior and memory pressure

Nice to have

  • Experience with Snapdragon Automotive, ARM Ethos, or specialized NPU pipelines
  • Background in running quantized LLMs on-device using llama.cpp or TFLite transformers
  • Familiarity with functional safety concepts (ISO 26262), sandboxing, or policy enforcement
  • Experience bridging cloud-trained models to resource-constrained embedded runtimes

What the JD emphasized

  • shipping ML inference on embedded, mobile, or automotive platforms
  • Expertise in model optimization techniques such as quantization, pruning, and compilation
  • Experience integrating LLM function calling or tool execution with structured outputs
  • Deep understanding of edge constraints including real-time behavior and memory pressure

Other signals

  • on-device ML systems
  • production environments
  • real-world constraints
  • latency, thermal limits, functional safety
  • end-to-end lifecycle