Data Scientist, Gtm Intelligence

OpenAI OpenAI · AI Frontier · San Francisco, CA · Data Science

Data Scientist role focused on building and operating intelligence systems for customer-facing teams. This involves defining and shipping production workflows, decision products, and agent workflows that improve account prioritization, identify risks/opportunities, and measure outcomes. The role requires a blend of technical depth in SQL and Python, product judgment, and stakeholder management, with a focus on shipping reliable, monitored decision products.

What you'd actually do

  1. Set the roadmap and methodology for GTM intelligence and decision products, using deep stakeholder discovery to probe beyond stated requests, uncover the underlying decisions, workflows, constraints, and measures of success, and translate them into measurable systems.
  2. Own the full lifecycle of intelligence products, including feature definition, methodology, evaluation, SQL and Python pipelines, scheduled refresh, serving, versioning, monitoring, and history.
  3. Build canonical feature datasets across product telemetry, commercial systems, CRM data, customer context, and field activity.
  4. Choose appropriately among heuristics, weighted scores, statistical models, ranking approaches, and machine-learning methods based on the decision, data maturity, and operational constraints.
  5. Partner closely with Technical Success and other GTM stakeholders as design partners: digging into their workflows, testing assumptions, and shaping the right solution to improve account prioritization, identify risks and opportunities, select interventions, and measure outcomes.

Skills

Required

  • Advanced SQL
  • Production Python
  • Feature engineering
  • Pragmatic model selection
  • Evaluation design
  • Calibration or threshold setting
  • Ongoing system monitoring
  • Building or owning reliable data transformations, canonical datasets, scheduled workflows, and application-facing outputs
  • Stakeholder discovery and communication skills

Nice to have

  • Databricks
  • Spark
  • dbt
  • Airflow
  • Modern cloud warehouses or lakehouses
  • B2B SaaS
  • Usage-based products
  • CRM or Salesforce data
  • Customer lifecycle systems
  • Recommendations
  • Next-best-action products
  • Model and feature versioning
  • Scheduled scoring
  • Reproducibility
  • Safe rollout
  • Defining exposure, action, feedback, and outcome data for decision products, experimentation, or impact measurement
  • Agentic systems
  • Data interfaces designed for both human and machine consumption

What the JD emphasized

  • own a flexible portfolio of high-impact decision data products
  • own the full lifecycle of intelligence products
  • personally ship reliable first versions
  • shipped and operated model-backed or rules-based decision products
  • move from source data through a production decision product
  • direct ownership of production decision systems
  • taking a score, signal, recommendation, ranking model, or decision rule from prototype into monitored production use
  • building or owning reliable data transformations, canonical datasets, scheduled workflows, and application-facing outputs

Other signals

  • Builds intelligence systems for customer-facing teams
  • Owns full lifecycle of intelligence products
  • Ships reliable production workflows and decision products
  • Partners with GTM stakeholders to improve account prioritization, identify risks and opportunities, select interventions, and measure outcomes