Staff Frontier Agents Engineer (applied Ai)

Scale AI Scale AI · Data AI · New York, NY +1 · Enterprise Engineering

Scale AI is seeking a Staff Frontier Agents Engineer to design, evaluate, and deploy intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. This role involves building production AI agents, architecting intelligent systems, developing multi-agent systems, and owning the experimentation lifecycle with a focus on production quality, reliability, and customer collaboration.

What you'd actually do

  1. Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use.
  2. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
  3. Own the full experimentation lifecycle, from hypothesis generation to production rollout.
  4. Design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
  5. Build production-quality AI systems with a strong emphasis on reliability, observability, latency, safety, and cost.

Skills

Required

  • 8+ years of software engineering, machine learning, or applied AI experience.
  • Strong Python programming skills.
  • Experience building production AI systems using LLMs.
  • Experience with modern AI tooling, including OpenAI, Claude, MCP, agent frameworks, vector databases, or retrieval systems.
  • Strong understanding of machine learning fundamentals and modern language models.
  • Experience designing or evaluating AI systems using quantitative metrics.
  • Excellent communication skills and the ability to work directly with enterprise customers.

Nice to have

  • Experience building production AI agents or autonomous systems.
  • Deep understanding of reasoning, retrieval, memory, planning, and tool use.
  • Experience designing evaluation frameworks for LLMs and agentic systems

What the JD emphasized

  • production AI systems
  • enterprise customers
  • production AI agents
  • production workflows
  • production systems
  • production AI engineering
  • production environments
  • production AI architectures
  • scalable production systems

Other signals

  • building production AI agents
  • customer intelligence platform
  • healthcare copilot
  • autonomous workflow
  • designing evaluation frameworks
  • production AI engineering
  • agent guardrails
  • human-in-the-loop workflows
  • partner directly with enterprise customers
  • translate ambiguous customer problems into production AI architectures
  • rapidly prototype new ideas
  • evolve successful solutions into scalable production systems
  • identify reusable patterns