Data AI · Open-source model infra
Currently tracking 22 active AI roles, down 17% versus the prior 4 weeks. Primary focus: Serve · Engineering. Salary range $121k–$300k (avg $227k).
Together AI currently has 25 active AI-related job listings. The majority of these roles, 80%, are focused on serving infrastructure. Engineering is the dominant function with 21 listings, and the United States is the primary hiring country with 19 roles. Frequent tech tags include model serving, inference infrastructure, and fine-tuning. Over the last 30 days, Together AI posted 4 new AI roles, representing a 33% decrease compared to the previous 30-day period.
Together AI currently has 26 active AI-related roles in our index. The most common open titles are: Solutions Architect (2), AI Infrastructure Engineer, AI Researcher, Core ML (Turbo), AI infrastructure Engineer (SRE) Bangalore , Customer Support Engineer (Inference), India. Most positions are in Engineering and Research.
Together AI's active AI hiring is concentrated in: serving infrastructure (81%), post-training (12%), application (4%). These categories follow a seven-stage AI lifecycle: data, pre-training, post-training, serving infrastructure, agents, evaluation, and application.
Together AI is hiring AI talent in: United States (20 roles), Netherlands (3 roles), United Kingdom (1 role).
Job postings at Together AI most frequently mention: Production ML Systems, Inference Infrastructure, GPU Computing, LLM Inference, Text-to-Speech.
In the past 30 days, Together AI has posted 5 new AI-related roles.
| Title | Stage | AI score |
|---|---|---|
| Research Engineer, Core ML Research Engineer role focused on improving inference efficiency and unifying it with RL/post-training systems for production-grade AI APIs. The role involves end-to-end ownership of critical systems, translating frontier ideas into robust infrastructure, and shipping measurable improvements in latency, throughput, cost, and model quality at scale. | ServePost-train | 10 |
| Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking Forward Deployed Engineer (FDE) focused on Inference & Post-Training for strategic customers. Responsibilities include optimizing inference engines, tuning performance for POCs and deployments, and guiding customers through fine-tuning pipelines (LoRA, SFT, DPO, RLHF, GRPO). Acts as a technical point of contact for strategic accounts, ensuring successful platform adoption and influencing product roadmap. | ServePost-train | 9 |
| Systems Research Engineer Intern - GPU Programming (Fall 2026) Internship role focused on optimizing GPU-accelerated kernels and algorithms for ML/AI applications, co-designing GPU kernels and model architecture with modeling teams, and contributing to efficient GPU architectures and programming models. | Serve | 9 |
| Staff Machine Learning Engineer, Voice AI Staff ML Engineer focused on optimizing the model serving layer for voice AI applications, including speech-to-text and text-to-speech models, with a focus on latency, throughput, and GPU utilization using inference engines like TRT-LLM and SGLang. The role involves building evaluation frameworks, supporting model partners, and shaping the architecture for next-generation voice models. | Serve | 9 |
| Senior Machine Learning Engineer, Voice AI Senior ML Engineer focused on optimizing the model serving layer for voice AI workloads, including speech-to-text and text-to-speech models. The role involves hands-on work with inference engines, GPU optimization, batching strategies, and ensuring new model architectures can be productionized efficiently. The goal is to achieve best-in-class latency and reliability for real-time voice applications. | Serve | 9 |
| Systems Research Engineer, GPU Programming This role focuses on optimizing and developing GPU-accelerated kernels and algorithms for ML/AI applications, requiring expertise in GPU programming (CUDA, Triton) and performance profiling. The engineer will collaborate with modeling, hardware, and software teams to enhance AI system efficiency and co-design GPU architectures. | Serve | 9 |
| AI Researcher, Core ML (Turbo) AI Researcher focused on the intersection of efficient inference algorithms, architectures, engines, and post-training/RL systems for production-scale API services. The role involves advancing inference efficiency, unifying inference with RL/post-training, and owning critical systems. | ServePost-train | 9 |
| Technical Support Engineer (Inference) - US Weekends Technical Support Engineer role focused on supporting customers using Together AI's inference and fine-tuning services, acting as a customer-facing SRE for inference endpoints on Kubernetes. Responsibilities include resolving technical challenges, ensuring endpoint health, managing incidents, contributing infrastructure changes for model deployment, and collaborating with engineering and product teams. Requires experience with Kubernetes, AI/ML/GPU technologies, infrastructure services, observability tooling, and LLM inference frameworks. | Serve | 8 |
| Staff Engineer, Distributed Storage and HPC & AI Infrastructure Staff Engineer focused on designing and delivering multi-petabyte storage systems optimized for AI training and inference workloads. Responsibilities include architecting high-performance parallel filesystems and object stores, building Kubernetes-native storage operators, optimizing data paths for high throughput, and implementing intelligent caching and data distribution strategies. The role requires deep expertise in distributed storage systems, Kubernetes, and programming in Go and Python. | Serve | 8 |
| LLM Inference Frameworks and Optimization Engineer Seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines for multimodal and language models. Focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for efficient large-scale AI deployment. | Serve | 8 |
| Machine Learning Engineer - Inference Machine Learning Engineer focused on optimizing and enhancing the performance of AI inference systems, working with state-of-the-art large language models to ensure efficient and effective operation at scale. Responsibilities include designing and building production systems, optimizing runtime inference services, and creating supporting tools and documentation. | Serve | 8 |
| AI infrastructure Engineer (SRE) Bangalore AI Infrastructure Engineer (SRE) responsible for the availability, reliability, and scalability of user-facing services and production systems, specifically focusing on GPU-enabled Kubernetes clusters for ML training and inference. | Serve | 7 |
| Lead/Manager Together Cloud Infrastructure Lead/Manager for Together Cloud Infrastructure in Amsterdam, focusing on building and managing a team to develop and operate a global, high-performance cloud platform for AI workloads, including GPU scheduling, management plane, and customer-facing services. | Serve | 7 |
| AI Infrastructure Engineer AI Infrastructure Engineer responsible for keeping user-facing services and production systems running smoothly, specializing in systems, availability, reliability, and scalability, with interests in algorithms and distributed systems. Builds and runs infrastructure using Ansible, Terraform, and Kubernetes, and develops monitoring systems. | Serve | 7 |
| Together Cloud Infrastructure Engineer This role focuses on building and maintaining the AI cloud infrastructure, including services for hardware management, IaaS software layer for GPU data centers, high-performance object storage for pretraining, and advanced observability stacks. The engineer will work on the core Together AI platform, create services and tools, and develop testing frameworks for robustness and fault-tolerance. | ServeData | 7 |
| Solutions Architect Solutions Architect at Together AI, a research-driven AI company focused on lowering the cost of AI systems. This role involves working with customers to build Generative AI applications using open-source models, acting as a technical advisor, running demos and POCs, and collaborating with sales. Requires strong technical background in AI/ML, GPU technologies, Python/JavaScript, and familiarity with infrastructure services. The role contributes to product feedback and educational content creation. | Serve | 7 |
| Senior Backend Engineer, Inference Platform Senior Backend Engineer focused on building and optimizing the inference platform for advanced generative AI models, including LLMs and multimodal models, at scale. The role involves optimizing latency, throughput, and resource allocation across tens of thousands of GPUs, collaborating with researchers to productionize frontier models, and contributing to open-source inference projects. | Serve | 7 |
| Platform Engineer, Model Shaping Platform Engineer focused on building and operating the foundational infrastructure for Together AI's model customization and evaluation platform. This includes backend services, scaling production workflows, and a job orchestration platform across datacenters and heterogeneous hardware. The role emphasizes reliability, CI/CD, and cloud/hybrid environment management. | ServeData | 7 |
| Senior Software Engineer - Together Cloud Infrastructure Senior Software Engineer focused on building and operating a high-performance, global AI cloud infrastructure platform. This includes designing and maintaining backend services for hardware management, IaaS software layer for GPU data centers, high-performance object storage for pretraining datasets, and advanced observability stacks for distributed pretraining. The role also involves architecture and research for decentralized AI workloads and contributing to the open-source platform. | ServeData | 7 |
| Solutions Architect Solutions Architect at Together AI to work with customers and prospects to create business value through Generative AI applications. This role involves acting as a technical advisor, running demonstrations and POCs, collaborating with sales, building relationships with customer leadership, delivering feedback to product/engineering/research, and building educational content. Requires 5+ years in a customer-facing technical role with 2+ years in pre-sales, strong technical background in AI/ML/GPU, understanding of LLM training/fine-tuning/inference, Python/JavaScript proficiency, and familiarity with infrastructure services. | Serve | 7 |