Member of Technical Staff - Model Serving / API Backend Engineer

Black Forest Labs Black Forest Labs · Multimodal · Freiburg, San Francisco · Engineering

This role focuses on bridging the gap between AI research breakthroughs and production systems by building and optimizing inference services and high-performance APIs for generative models. The engineer will own the productionization of research checkpoints, ensuring low latency, high throughput, and scalability for millions of requests, while also improving monitoring and observability.

What you'd actually do

  1. Turn research checkpoints into production-ready inference services
  2. Design and maintain high-performance APIs serving millions of requests
  3. Optimize inference latency and throughput across GPU infrastructure
  4. Build scalable serving architectures that handle unpredictable traffic
  5. Improve reliability, monitoring, and observability across model-serving systems

Skills

Required

  • Python
  • FastAPI
  • async systems
  • GPU infrastructure
  • CUDA
  • inference optimization
  • Docker
  • Kubernetes
  • Redis
  • Postgres
  • distributed task queues
  • Cloud platforms (AWS, GCP, or Azure)
  • Observability stacks (metrics, logging, tracing)
  • system design
  • debugging
  • deployment
  • performance optimization
  • reliability engineering
  • cost optimization

Nice to have

  • Real-time or low-latency inference systems
  • TensorRT
  • reduced precision
  • layer fusion
  • model compilation techniques
  • Frontend demo tooling (Streamlit, Gradio, React)
  • CI/CD and automated testing for ML systems
  • Security best practices for API and model serving

What the JD emphasized

  • productionization
  • inference latency
  • API serving
  • GPU infrastructure
  • scaling APIs
  • ML systems under load
  • production ML serving
  • production ML systems

Other signals

  • productionization
  • inference optimization
  • API serving
  • GPU infrastructure