Sr. Inference Optimization Engineer (local / Edge Runtime)

Intel Intel · Semiconductors · California, Santa Clara, United States +3

This role focuses on optimizing AI inference engines (like llama.cpp, vLLM) for local and edge hardware, aiming to improve latency, throughput, and memory usage. The engineer will tune key inference parameters, drive quantization strategies, reduce CPU overhead, and benchmark performance across different hardware tiers, with a goal of making hybrid, low-cost agent products viable.

What you'd actually do

  1. Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware
  2. Tune KV cache, continuous batching, and scheduling for interactive agent workloads
  3. Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team
  4. Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)
  5. Benchmark across hardware tiers and publish honest performance comparisons

Skills

Required

  • BS/MS in CS, EE, Math or related STEM field
  • 8+ years software development background
  • Strong in C++ and/or Python
  • Experience with LLM inference (attention, KV cache, decoding)
  • Experience profiling and optimizing real performance problems (CPU or GPU)
  • Linux, build systems, and low-level debugging expertise

Nice to have

  • Hands-on with llama.cpp, vLLM, ggml, or similar engines
  • Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD / CPU kernels
  • Familiarity with quantization formats and their quality trade-offs
  • Open-source contributions to inference engines

What the JD emphasized

  • local inference
  • edge hardware
  • latency
  • throughput
  • memory
  • KV cache
  • batching
  • scheduling
  • quantization
  • CPU overhead
  • engine startup
  • model load
  • lifecycle
  • performance comparisons
  • local / edge AI

Other signals

  • Optimizing inference engines for local and edge environments
  • Focus on latency, throughput, and memory on constrained hardware
  • Tuning KV cache, batching, and scheduling for agent workloads
  • Driving quantization strategy and validating quality impact
  • Reducing CPU overhead and improving engine startup/lifecycle