Applied AI Engineer, Gtm Growth Engineering

OpenAI OpenAI · AI Frontier · San Francisco, CA · Go To Market

Applied AI Engineer to build production systems for AI-powered go-to-market workflows, focusing on the agent improvement loop (behavior, feedback, evaluation, experimentation) to enhance effectiveness and reliability. This role involves end-to-end ownership, instrumenting workflows, defining quality standards, investigating performance issues, designing improvements, building backend services, running experiments, and partnering with various teams to drive business outcomes with safeguards.

What you'd actually do

  1. Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes.
  2. Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context.
  3. Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring for real GTM workflows.
  4. Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design.
  5. Design and ship targeted behavior improvements, including changes to prompting, context construction, decision logic, tool use, and human-review paths.

Skills

Required

  • 4+ years of software, backend, applied AI, or product-engineering experience building reliable production systems.
  • Experience building AI agents, LLM-powered applications, or other model-driven workflows that operated on real production traffic.
  • Experience diagnosing and improving agent behavior using production traces, user feedback, evaluation, experimentation, or careful systems design.
  • Practical experience with evaluation design, regression testing, human or model grading, online quality signals, or controlled experiments.
  • Strong backend engineering skills across Python, APIs, data pipelines, stateful workflows, and production services.
  • Strong product judgment and the ability to connect technical changes to customer experience, conversion, qualified pipeline, or operational efficiency.
  • Comfort working across model behavior, context, knowledge, tools, workflow state, and human-in-the-loop decisions.
  • The ability to work closely with technical and non-technical partners across Engineering, Product, Data Science, Sales, and B2B Marketing.
  • A pragmatic mindset: you can scope ambiguous problems, ship useful improvements, and build toward a durable system.

Nice to have

  • Experience building agent evaluation, observability, experimentation, or AI infrastructure products.
  • Experience with production replay, LLM grading, human-labeled datasets, shadow evaluation, or staged rollout.
  • Experience improving model or agent behavior through context design, prompting, tools, decision logic, or feedback loops.
  • Experience with sales, B2B marketing, revenue, CRM, campaign, or other GTM-facing systems.
  • Experience measuring customer engagement, qualified pipeline, conversion, or operational efficiency.

What the JD emphasized

  • build production systems
  • operated on real production traffic
  • diagnosing and improving agent behavior
  • evaluation design
  • backend engineering skills
  • product judgment
  • improve from real usage instead of stopping at a successful prototype
  • tracing messy production failures
  • evaluation is valuable when it helps teams make better product decisions
  • trustworthy deployment

Other signals

  • build production systems
  • agent improvement loop
  • end-to-end ownership
  • real-world signals
  • measurable improvements