AI Infrastructure Engineer

Intel Intel · Semiconductors · California, Santa Clara, United States +3

AI Infrastructure Engineer focused on optimizing LLM inference performance on Intel GPUs, involving kernel development, integration with serving frameworks (vLLM, SGLang), and influencing future hardware roadmaps.

What you'd actually do

  1. Drive Inference Performance: Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs.
  2. Deep Stack Optimization: Profile, diagnose, and resolve cross-stack performance bottlenecks.
  3. Kernel Development and Integration: Design, write, and optimize custom high-performance kernels for critical attention mechanisms, MoE, quantization, and operator fusions.
  4. Open Source Leadership: Upstream your architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch, acting as a bridge between the hardware teams and the open-source community.
  5. Shape the Hardware Roadmap: Apply roofline analysis and systematic profiling to decompose bottlenecks. You will partner with our architecture and compiler teams to shape future GPU roadmaps based on real-world GenAI workload data.

Skills

Required

  • modern C++
  • Python
  • GPU computing
  • AI systems
  • high-performance computing (HPC)

Nice to have

  • CPU/GPU architecture understanding
  • modern LLM architectures and inference paradigms
  • open-source contributions to inference engines
  • custom GPU kernels using Triton, SYCL, CUDA/CUTLASS
  • scale-out inference orchestration

What the JD emphasized

  • performance-obsessed
  • push LLM inference to its absolute limits
  • redefine peak performance
  • custom high-performance kernels
  • architectural improvements
  • real-world GenAI workload data

Other signals

  • LLM inference performance optimization
  • GPU kernel development
  • open-source contributions to inference frameworks
  • shaping future hardware roadmaps based on AI workload data