Director, AI & Tooling

Salesforce Salesforce · Enterprise · Bellevue, Illinois - Chicago, Massachusetts - Boston, Texas - Dallas, California - Irvine, Georgia - Atlanta, Indiana - Indianapolis, Texas - Austin, WA

Salesforce is seeking a Director of AI & Tooling to lead a team focused on generative AI capabilities, evaluation tooling, and observability for their customer-facing Help Agent. This role involves driving AI product strategy, data-driven evaluation, and production excellence to ensure the Help Agent performs reliably and delivers business value. Responsibilities include leading generative AI strategy, owning evaluation tooling, managing observability programs, overseeing the NNAOV program, and supporting AI skill development. The goal is to improve resolution rates, quality scores, and latency, and establish a centralized observability platform and scalable evaluation framework.

What you'd actually do

  1. Serve as the AI & Tooling team lead within Digital Success Data & AI, setting team mission, operating model, and priorities
  2. Act as a Customer Zero steward — representing Salesforce's own use of Agentforce as a reference customer, surfacing product gaps and driving roadmap influence through Shiproom intakes and direct T&P partnerships
  3. Monitor and manage Help Agent production health, including end-to-end latency, engagement rate, escalation rate, resolution rate, and quality scores — escalating issues and coordinating cross-functional remediation as needed
  4. Lead cross-functional collaboration with Technology Teams to align on evaluations, retrieval strategies, and tooling development
  5. Define and mature the AI Evaluations Framework, moving from manual, spreadsheet-based processes toward scalable, automated evaluation at production scale

Skills

Required

  • Generative AI strategy
  • LLM configurations
  • prompt and planner optimization
  • agentic readiness scoring
  • evaluations tooling strategy
  • Testing Center
  • Agentforce Testing Agents (ATA)
  • custom evaluation frameworks
  • answer quality metrics (completeness, correctness, relevance, faithfulness, clarity, latency)
  • Help Agent Observability
  • real-time performance monitoring
  • session tracing
  • error analysis
  • failure clustering
  • business impact measurement
  • NNAOV Program Management
  • AI Skill & Application Development
  • Slackbot
  • Claude Code
  • production health monitoring
  • end-to-end latency
  • engagement rate
  • escalation rate
  • resolution rate
  • quality scores
  • cross-functional collaboration
  • AI Evaluations Framework
  • scalable, automated evaluation
  • production-grade monitoring platform
  • leadership forums
  • senior stakeholders

Nice to have

  • Ensemble Retriever

What the JD emphasized

  • evaluation tooling
  • evaluations tooling
  • evaluation frameworks
  • evaluations
  • evaluation
  • observability
  • Help Agent
  • Agentforce
  • Customer Zero

Other signals

  • AI CRM
  • generative AI capabilities
  • evaluation tooling
  • observability
  • Customer Zero implementations
  • Help Agent
  • Agentforce