Senior AI Engineer - Services Special Projects

Apple Apple · Big Tech · San Francisco Bay Area, CA +1 · Software and Services

Senior AI Engineer focused on building, deploying, optimizing, and operationalizing LLM-based applications with a strong emphasis on MLOps/LLMOps and scalable production systems. The role involves owning the infrastructure and tooling for LLM-powered features, including CI/CD pipelines, serving infrastructure, observability, versioning, and governance. It covers the full model lifecycle from experimentation and fine-tuning to deployment, monitoring, and retirement, with a focus on multimodal data, cloud-native infrastructure, model optimization (quantization, distillation), and safety guardrails. The role also includes mentoring and setting technical standards.

What you'd actually do

  1. Own the full model lifecycle: from experimentation and training through validation, deployment, monitoring, and retirement, ensuring reproducibility and governance at every stage.
  2. Fine-tune and tune models, including hyperparameters, adapters/LoRA, and distillation targets, to improve quality, task fit, and efficiency.
  3. Design and build scalable ML infrastructure and experimentation platforms, including web-based interfaces, dashboards, and backend services, that enable rapid model development, testing, and deployment at scale.
  4. Define and implement CI/CD methodologies for model integration, deployment, versioning, and monitoring, and build the production infrastructure, including cloud-native deployment (Kubernetes, AWS) and well-modeled RESTful/GraphQL APIs, that serves high-traffic LLM services reliably and cost-efficiently.
  5. Optimize models for production, including quantization, distillation, and compilation (e.g., ONNX, TensorRT), tuning for token throughput, latency, and cost targets.

Skills

Required

  • Master's degree in Computer Science, Engineering, or a related field
  • 8+ years of experience in Machine learning and software engineering
  • Proven track record of shipping production-grade ML/LLM systems
  • Strong understanding of LLMs, fine-tuning, prompt engineering, and RAG patterns
  • Experience building pipelines that process multimodal data (structured and image) and integrate ML model inference, including LLMs and embedding models, for data enrichment and transformation
  • Hands-on experience deploying, serving, and optimizing LLMs or ML models in production, including inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), serving frameworks (Triton, vLLM, SGLang, TorchServe, or similar), and tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput inference
  • Experience with vector search technologies (e.g., Pinecone, Milvus) and storing/serving embeddings (e.g., pgvector, FAISS)
  • Experience with feature stores (e.g., Feast) and data lineage tracking
  • Strong proficiency in Python, with solid software engineering fundamentals, including backend service frameworks (e.g., Flask, FastAPI), for building ML/LLM services, pipelines, and tooling
  • Working proficiency in Java or Scala, sufficient to integrate with JVM-based data infrastructure (e.g., Spark, Flink, Kafka clients) and the broader services platform.
  • Experience with distributed systems, cloud platforms (e.g., AWS), container orchestration (Kubernetes), CI/CD pipelines, and building Data Pipelines on Spark using Airflow
  • Experience with ML lifecycle management and versioning practices, including experiment tracking, model registry, deployment automation, and dataset/model versioning tools (e.g., DVC, MLflow, Weights & Biases, Delta Lake)
  • Experience with workflow orchestration platforms (Airflow)
  • Excellent communication

What the JD emphasized

  • shipping production-grade ML/LLM systems
  • LLMs
  • fine-tuning
  • prompt engineering
  • RAG patterns
  • multimodal data
  • ML model inference
  • LLMs
  • embedding models
  • deploying, serving, and optimizing LLMs or ML models in production
  • inference runtimes/compilers
  • serving frameworks
  • tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput inference
  • vector search technologies
  • storing/serving embeddings
  • feature stores
  • data lineage tracking
  • Python
  • backend service frameworks
  • ML/LLM services, pipelines, and tooling
  • Java or Scala
  • JVM-based data infrastructure
  • distributed systems
  • cloud platforms
  • container orchestration
  • CI/CD pipelines
  • Data Pipelines on Spark
  • ML lifecycle management
  • versioning practices
  • experiment tracking
  • model registry
  • deployment automation
  • dataset/model versioning tools
  • workflow orchestration platforms

Other signals

  • MLOps/LLMOps
  • scalable production systems
  • full model lifecycle
  • ML infrastructure
  • model governance
  • privacy as a design constraint