Applied AI Engineer

Intel Intel · Semiconductors · Bangalore, India

Intel is seeking a senior contributor for their Data Center and AI group (DCAI) to work on open-source AI frameworks like PyTorch, Tensorflow, and JAX. The role involves designing and developing features for Intel's AI frameworks software stack, enabling and optimizing software for Intel's AI accelerators and GPUs, and identifying performance optimization opportunities for deep learning workloads. Responsibilities include ML kernel development, enhancing training and inference capabilities, and participating in the open-source community.

What you'd actually do

  1. Design and develop SW features for AI frameworks - both HW-agnostic and HW-aware, especially in ML kernel development.
  2. Enhance and extend the Deep learning training, and Inference capabilities in the Software stack.
  3. Identifying optimization opportunities in the software stack to enhance performance of Deep learning workloads
  4. Participate in discussions with Open-source community, involve in development, adopting upstream and Upstream software.

Skills

Required

  • Advanced C++ (C++ 14/17)
  • Python
  • Parallel programming
  • Machine learning kernels (GEMM, Convolution, Flash attention)
  • PyTorch, Tensorflow, or JAX
  • Deep Learning models/LLMs for Vision / NLP
  • Debugging complex issues in multi-layered SW systems
  • SW integration in large open-source frameworks
  • Computer architecture
  • HW-SW optimization techniques

Nice to have

  • Developing and integrating CUTLASS or Triton based kernels in Large language models (LLMs)
  • Compiler algorithms for heterogeneous systems
  • Fuser optimizations

What the JD emphasized

  • Proficient in Advanced C++ (C++ 14/17)
  • Experience in developing machine learning kernels such as GEMM, Convolution, Flash attention etc
  • In depth and hands on experience in one of the frameworks such as PyTorch, Tensorflow or JAX
  • Practical knowledge of Deep Learning models/LLMs for Vision / NLP
  • Understanding of SW integration in large open-source frameworks.
  • Strong understanding of computer architecture and HW-SW optimization techniques.
  • Experience in working on frameworks/platforms that have gone to production.

Other signals

  • optimizing state of the art Software stack for Intel's AI accelerators
  • enabling and optimizing state of the art Software stack for Intel's AI accelerators and next generation GPUs
  • ML kernel development
  • Deep learning training, and Inference capabilities
  • optimizing opportunities in the software stack to enhance performance of Deep learning workloads