Technical Support Engineer (gpu Clusters) - US Weekends

Together AI Together AI · Data AI · San Francisco, CA · Customer Success

This role is a customer-facing Technical Support Engineer focused on supporting customers using Together AI's GPU clusters for training, fine-tuning, and inference. The engineer will act as a product expert and SRE, troubleshooting complex technical challenges related to Kubernetes, GPU hardware, networking, and storage. They will collaborate with engineering and product teams to drive improvements and transform customer insights into product roadmap actions. The role requires experience with AI/ML infrastructure, Kubernetes, and high-performance computing environments.

What you'd actually do

  1. Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge Kubernetes GPU clusters; ensure swift and effective solutions every time.
  2. Act as a customer facing SRE to ensure our customer’s Kubernetes clusters remain healthy and stable
  3. Become a product expert in our GPU Cluster service, serving as the last line of technical defense before issues are escalated to Engineering and Product teams.
  4. Monitor GPU cluster health and proactively communicate hardware issues to customers (thermal throttling, BMC failures, NVLink/InfiniBand degradation) with clear remediation steps
  5. Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management

Skills

Required

  • 3+ years of experience in a customer-facing technical role with at least 1 year in a support function for an AI service or supporting a mission-critical API in SaaS
  • Experience as an SRE or DevOps engineer working with Kubernetes
  • Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments.
  • Advanced knowledge with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting and programming languages.
  • Experience with HPC/Slurm cluster environments — node draining, job scheduling, maintenance workflows
  • Familiarity with high-speed networking concepts — InfiniBand, RDMA, network interface diagnostics
  • Experience with distributed storage systems (e.g., Weka, NFS) and troubleshooting I/O and bandwidth issues
  • Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters.
  • Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution
  • Ability to work cross-functionally with teams such as Sales, Engineering, Support, Product and Research to drive customer success.
  • Strong sense of ownership and willingness to learn new skills to ensure both team and customer success.
  • Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders.
  • Ability to operate in dynamic environments, adept at managing multiple projects, and comfortable with frequent context switching and prioritization.

What the JD emphasized

  • customer-facing technical role
  • support function for an AI service
  • Kubernetes
  • AI, ML, GPU technologies
  • high-performance computing (HPC) environments
  • Kubernetes
  • SLURM
  • high-performance network fabrics
  • distributed storage systems
  • troubleshooting I/O and bandwidth issues
  • compute clusters