Applied AI Inference Engineer

Crusoe Crusoe · Data AI · San Francisco, CA - US · Cloud Engineering

Applied AI Inference Engineer focused on optimizing large language models for speed, cost, and reliability in production. This role involves end-to-end ownership of the inference stack, from profiling and optimization to customer tailoring and monitoring. Requires hands-on coding, profiling, and low-level optimization skills, with a customer-facing component.

What you'd actually do

  1. Bring current inference techniques into production and refine them.
  2. Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
  3. Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems.
  4. Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models.
  5. Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic.

Skills

Required

  • shipping code in production
  • Python
  • optimizing LLMs for high throughput / low latency inference
  • modern LLM serving frameworks such as vLLM or SGLang
  • profiling and analyzing performance down to the kernel level
  • how GPUs are built and how they behave
  • large language models
  • AI/ML pipelines and the full path of developing and deploying ML models
  • communication skills

Nice to have

  • making software systems run faster, especially for large language models
  • CUDA
  • software engineering fundamentals
  • building and shipping AI/ML inference systems
  • Docker
  • Kubernetes
  • building or tuning AI/ML projects, particularly in a customer-facing setting

What the JD emphasized

  • core systems and performance work
  • applied, not academic
  • hands-on engineering role built around coding, profiling, and low-level optimization
  • customer-facing side
  • performance work gets proven

Other signals

  • LLM inference optimization
  • production deployments
  • customer-facing engineering