Machine Learning/ Search Engineer - Services Special Projects

Apple Apple · Big Tech · Cupertino, CA +1 · Software and Services

This role focuses on building and optimizing large-scale, real-time search systems at Apple, integrating Generative AI and Information Retrieval. The engineer will develop ML/AI models for ranking and relevance, work with embedding models and vector databases, and deploy/optimize LLMs and embeddings in the production query path. The role requires strong experience in search infrastructure, ML frameworks, and scalable systems.

What you'd actually do

  1. Design, build, and maintain large-scale, low-latency, high-performance search systems that can scale.
  2. Develop and optimize ranking, relevance, and retrieval through ML/AI models and merging traditional keyword search with vector-based semantic search using embedding models and vector databases.
  3. Develop sophisticated NLP pipelines for intent classification, entity extraction, semantic parsing, and query expansion.
  4. Merge traditional keyword search (BM25) with vector-based semantic search using embedding models and vector databases.
  5. Design and Implement machine learning models (e.g. Learning to Rank, Cross Encoder based models) and multi-stage reranking algorithms to optimize search precision and recall.

Skills

Required

  • Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or a related field
  • 10+ years of experience in Machine Learning, Data Science, or Software Engineering roles with a significant focus on search infrastructure and information retrieval.
  • Hands on experience building and deploying large-scale search systems in production.
  • Deep understanding of information retrieval, query understanding, query augmentation and multi-stage ranking algorithms
  • Strong foundation in deep learning architectures for search and retrieval (e.g., transformers, cross encoder models, graph neural networks, learned sparse representations).
  • Experience with to multi-objective optimization in search systems (e.g., relevance, diversity, freshness, fairness).
  • Experience with real-time systems, user feedback loops, and model retraining pipelines.
  • Strong proficiency in Go, Java, C++ and Python
  • Proven experience with ML frameworks including PyTorch, XGBoost.
  • Familiarity with cloud environments (including AWS) and containerization (Docker, Kubernetes)
  • Extensive experience working with data processing pipelines including Spark, Flink
  • Hands-on experience with vector search including FAISS
  • Familiarity with streaming platforms including Apache Kafka
  • Experience with search infrastructure including OpenSearch, and/or Elasticsearch
  • Hands-on experience deploying, serving, and optimizing LLMs, Embeddings and ML models directly in the production query/request path
  • Past successful deployments with tuning of models (including quantization) for performance and quality optimization
  • Excellent communication skills and a collaborative mindset

Nice to have

  • Master's Degree; PhD Preferred
  • Published work or patents in the domain of search systems, information retrieval, or related ML fields.
  • Experience with graph databases such as TigerGraph
  • Experience with data and model versioning tools and practices (e.g., DVC, MLflow, Weights & Biases)
  • Deep Experience with KV Stores including SSTables and Cassandra
  • Experience with tuning KV-cache and batching for low-latency, high-throughput real-time inference.
  • Deep production level experience with inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (vLLM, SGLang or Triton, TorchServe ) .

What the JD emphasized

  • 10+ years of experience in Machine Learning, Data Science, or Software Engineering roles with a significant focus on search infrastructure and information retrieval.
  • Hands on experience building and deploying large-scale search systems in production.
  • Hands-on experience deploying, serving, and optimizing LLMs, Embeddings and ML models directly in the production query/request path
  • Past successful deployments with tuning of models (including quantization) for performance and quality optimization

Other signals

  • building a massive, real-time search experience from the ground up
  • search at the intersection of Generative AI and Information Retrieval
  • crafting intelligent systems that personalize user experiences
  • Develop and optimize ranking, relevance, and retrieval through ML/AI models
  • Hands-on experience deploying, serving, and optimizing LLMs, Embeddings and ML models directly in the production query/request path