Backend Engineer - API

xAI xAI · AI Frontier · Palo Alto, CA · Engineering

Backend Engineer responsible for building and owning the xAI API, focusing on high-throughput, low-latency inference for LLMs. This includes model serving infrastructure, request routing, rate limiting, observability, and scaling, with potential involvement in agent SDKs and orchestration.

What you'd actually do

  1. Build the xAI API that serves our models to developers worldwide
  2. Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling

Skills

Required

  • Rust
  • C++
  • distributed systems
  • observability
  • reliability
  • PostgreSQL
  • Clickhouse
  • MongoDB
  • gRPC

Nice to have

  • LLM inference engines
  • serving frameworks
  • agent SDKs
  • agent orchestration frameworks
  • Docker
  • Kubernetes
  • containerized applications

What the JD emphasized

  • Expert knowledge of either Rust or C++
  • Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems
  • Knowledge of service observability and reliability best practices
  • Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM)
  • Experience designing or building with agent SDKs and agent orchestration frameworks

Other signals

  • high-throughput inference
  • low latency
  • high availability
  • model serving infrastructure
  • request routing
  • rate limiting
  • observability
  • efficient scaling
  • LLM inference engines
  • serving frameworks
  • agent SDKs
  • agent orchestration frameworks