Staff Forward Deployed Engineer

Tenstorrent Tenstorrent · Semiconductors · NORTH AMERICA · Product

Staff Forward Deployed Engineer role focused on bridging customer needs with Tenstorrent's AI inference hardware and software. Responsibilities include contributing production code, operating deployments, debugging the full inference stack, and providing customer feedback through pull requests and benchmarks. Requires strong software engineering skills, Kubernetes experience, and familiarity with LLM inference serving engines.

What you'd actually do

  1. create continuity between customers, engineering, and AI inference service products
  2. contribute production code, operate deployments, and you can explain a trade-off to customer leadership as clearly as to core engineering teams
  3. understand how accelerator compute, memory, and networking topology constrain AI workloads, and don't treat hardware as a black box
  4. work directly with customers to understand their challenges and provide effective solutions
  5. comfortable debugging across the full inference stack: from failing requests, through the serving layer, down to OOMs or kernel dispatch if need be

Skills

Required

  • 5+ years of relevant technical experience (e.g. Applied Engineer, Machine Learning Engineer, MLOps Engineer, Platform Engineer, Infrastructure Engineer, Site Reliability Engineer, Field Application Engineer)
  • Experience turning ambiguous customer requirements or issues into verifiable acceptance criteria
  • Kubernetes and Helm experience at multi-node, HPC, or AI cluster scale
  • Experience with observability and infrastructure automation, e.g. Prometheus, Grafana, OpenTelemetry
  • Experience with LLM inference serving engines and technologies, e.g. vLLM, SGLang, Mooncake, NIM, Dynamo, LMCache

Nice to have

  • early adopter of AI for your work from coding to building agentic workflows that multiply your impact

What the JD emphasized

  • production code
  • customer leadership
  • debugging across the full inference stack
  • LLM inference serving engines

Other signals

  • customer-facing role
  • production code
  • inference stack
  • LLM serving engines