Currently tracking 48 active AI roles, up 103% versus the prior 4 weeks. Primary focus: Data · Engineering. Salary range $148k–$600k (avg $302k).
xAI currently has 56 active AI-related job listings. The majority of these roles, 57%, are focused on data. Engineering is the top function for these positions, with the United States being the primary hiring country. Frequent tech tags include audio_speech, synthetic_data, and model_serving. Over the last 30 days, xAI has posted 11 new AI roles, representing a 267% increase compared to the previous 30-day period.
xAI currently has 73 active AI-related roles in our index. The most common open titles are: AI Tutor - Arabic, AI Tutor - Audio Editing, AI Tutor - Azerbaijani, AI Tutor - Bulgarian, AI Tutor - Catalan. Most positions are in Engineering.
xAI's active AI hiring is concentrated in: data (70%), application (8%), serving infrastructure (8%). These categories follow a seven-stage AI lifecycle: data, pre-training, post-training, serving infrastructure, agents, evaluation, and application.
xAI is hiring AI talent in: United States (43 roles), United Kingdom (2 roles), Japan (1 role).
Job postings at xAI most frequently mention: Speech & Audio Processing, Linguistics, Language Modeling, Data Labeling, Dialog Systems.
In the past 30 days, xAI has posted 30 new AI-related roles.
| Title | Stage | AI score |
|---|---|---|
| Member of Technical Staff - RL Inference Engineer to optimize inference stack for RL workloads, analyze performance bottlenecks, and implement novel RL techniques. Requires experience in distributed systems, LLM inference, and programming languages like Python/C++/Rust with frameworks like PyTorch/Jax/CUDA. Preferred skills include quantization, numerics, and inference engines. | ServePost-train | 9 |
| Member of Technical Staff - Inference The role focuses on designing and optimizing large-scale model serving systems for high-performance inference, ensuring speed and reliability for millions of users. Responsibilities include architecting distributed infrastructure, optimizing latency and throughput, building high-concurrency systems, and accelerating inference engines. | Serve |
| 9 |
| ML Infrastructure Engineer xAI is seeking an ML Infrastructure Engineer to build and optimize the reliable, high-performance ML platform powering their recommendations. This role involves designing and scaling GPU compute infrastructure, training frameworks, and experimentation tools, developing data pipelines, and integrating large-scale data, training, and inference systems. The engineer will collaborate with ML teams to productionize models, ensure scalability, reliability, and efficiency, and work across the full stack. Requires 2+ years of industry experience with large-scale production environments, distributed systems, GPU infrastructure, and ML platforms, proficiency in Python and C++/Rust, and familiarity with ML frameworks like JAX/PyTorch and Linux/orchestration tools. | Serve | 8 |
| Backend Engineer - API Backend Engineer responsible for building and owning the xAI API, focusing on high-throughput, low-latency inference for LLMs. This includes model serving infrastructure, request routing, rate limiting, observability, and scaling, with potential involvement in agent SDKs and orchestration. | ServeAgent | 8 |
| Software Engineer - Platform Infrastructure (Rust, C++) Software Engineer focused on building and optimizing large-scale distributed systems and infrastructure for AI training clusters, involving low-level systems programming and hardware/software co-design. | ServeData | 7 |
| Software Engineer: Network (C++) Develops core networking software for a massive GPU cluster powering AI models, focusing on high-performance, low-latency datacenter fabric for training efficiency and reliability. | Serve | 7 |
| Backend Engineer - API Backend Engineer responsible for building and owning the xAI API that serves models to developers worldwide. This includes end-to-end system ownership for high-throughput, low-latency, and highly available inference, covering model serving infrastructure, request routing, SDK development, rate limiting, observability, and scaling. Requires expertise in Rust or C++, distributed systems, and observability, with preferred experience in LLM inference engines, serving frameworks, and agent orchestration. | ServeAgent | 7 |
| Member of Technical Staff - Voice Product The role focuses on building and shipping voice features for the Grok product, involving backend engineering for scalable, low-latency voice infrastructure and model integrations. The goal is to drive performance, reliability, and quality of voice interactions at a global scale. | Serve | 7 |
| Member of Technical Staff - Compute Infrastructure The role involves building and optimizing large-scale GPU clusters and the platform layer for AI training and inference. Responsibilities include low-level CUDA kernel development, Linux kernel internals, custom orchestration, and performance debugging across the full stack to accelerate AI model development. | ServeAgent | 7 |
| Sr. Software Engineer (Data Center Automation) This role focuses on automating processes, building and implementing robust observability solutions, and ensuring seamless operations for mission-critical AI infrastructure within a multi-data center environment. The engineer will combine strong coding abilities with data center experience to build scalable reliability services, optimize system performance, and minimize downtime, with a focus on reducing MTTR and ensuring resilience for AI workloads. | Serve | 5 |
| Member of Technical Staff Seeking a Member of Technical Staff to manage and enhance reliability for a multi-data center AI infrastructure. This role focuses on automating processes, building observability solutions, and ensuring seamless operations for mission-critical AI infrastructure. The ideal candidate will combine strong coding abilities with hands-on data center experience to build scalable reliability services, optimize system performance, and minimize downtime, including partnership with facility operations. The primary objective is to mitigate downtime and minimize impact to end-users through proactive automation, robust observability, and integrated software-physical reliability strategies, ensuring AI infrastructure remains resilient and scalable. | Serve | 5 |