Senior Data Scientist – Applied AI

Salesforce Salesforce · Enterprise · Palo Alto, CA

Salesforce is seeking a Senior Data Scientist to join their Data Science team, focusing on building and evaluating enterprise AI systems including conversational agents, voice experiences, language models, and intelligent automation. The role involves designing experiments, building evaluation frameworks, analyzing telemetry, developing statistical models, optimizing prompts, fine-tuning foundation models, and ensuring responsible AI practices. The ideal candidate will have strong experience with LLMs, generative AI, AI agents, RAG, and fine-tuning, with a focus on delivering measurable customer impact.

What you'd actually do

  1. Design and execute experiments to evaluate and improve the quality of large language models (LLMs), voice/text AI systems, multimodal models, and long horizon task agents
  2. Build scalable evaluation datasets, benchmarks, and automated evaluation frameworks across language, speech, reasoning, and agent workflows
  3. Analyze large-scale product, customer, and model telemetry to identify failure modes, performance bottlenecks, and opportunities for improvement
  4. Develop statistical models, predictive analytics, and experimentation frameworks to measure model quality, user experience, and business impact
  5. Develop, optimize, and evaluate prompts, system instructions, retrieval strategies, and context engineering techniques to improve performance, reliability, efficiency, and safety of AI applications

Skills

Required

  • Python
  • Large Language Models (LLMs)
  • Generative AI
  • AI Agents
  • Prompt Engineering
  • Context Engineering
  • Foundation Model Fine Tuning
  • Reinforcement Learning (RL)
  • Retrieval-Augmented Generation (RAG)
  • Machine Learning
  • Statistical Modeling
  • Experiment Design and A/B Testing
  • Causal Inference
  • Predictive Analytics
  • AI Evaluation and Benchmarking
  • Responsible AI

Nice to have

  • Speech and Multimodal AI
  • Data Visualization

What the JD emphasized

  • evaluate and improve the quality of large language models (LLMs)
  • Build scalable evaluation datasets, benchmarks, and automated evaluation frameworks
  • Develop scalable evaluation methodologies for Responsible AI, including safety, robustness, fairness, security, and governance

Other signals

  • building AI systems at enterprise scale
  • evaluating and improving LLMs
  • responsible AI