Senior Frontier Agents Engineer (applied Ai)

Scale AI Scale AI · Data AI · New York, NY +1 · Enterprise Engineering

This role focuses on building and deploying production AI agents and intelligent systems for enterprise customers. It involves designing agent architectures, integrating LLMs with structured knowledge and traditional ML, developing multi-agent systems, and creating robust evaluation frameworks. The role emphasizes translating frontier AI research into reliable, scalable production systems that deliver measurable business impact.

What you'd actually do

  1. Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use.
  2. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
  3. Design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
  4. Build production-quality AI systems with a strong emphasis on reliability, observability, latency, safety, and cost.
  5. Partner directly with enterprise customers to understand their business, data, and operational challenges.

Skills

Required

  • 5+ years of software engineering, machine learning, or applied AI experience.
  • Strong Python programming skills.
  • Experience building production AI systems using LLMs.
  • Experience with modern AI tooling, including OpenAI, Claude, MCP, agent frameworks, vector databases, or retrieval systems.
  • Strong understanding of machine learning fundamentals and modern language models.
  • Experience designing or evaluating AI systems using quantitative metrics.
  • Excellent communication skills and the ability to work directly with enterprise customers.

Nice to have

  • Experience building production AI agents or autonomous systems.
  • Deep understanding of reasoning, retrieval, memory, planning, and tool use.
  • Experience designing evaluation frameworks for LLMs and agentic system

What the JD emphasized

  • production AI agents
  • intelligent systems
  • multi-agent systems
  • production workflows
  • enterprise data
  • production AI systems
  • agent guardrails
  • enterprise environments
  • production AI agents or autonomous systems
  • reasoning, retrieval, memory, planning, and tool use
  • evaluation frameworks for LLMs and agentic system

Other signals

  • building production AI agents
  • designing and deploying production AI agents
  • architect intelligent systems
  • develop multi-agent systems
  • translate frontier AI research into production systems
  • design rigorous evaluation frameworks
  • build production-quality AI systems
  • design agent guardrails
  • partner directly with enterprise customers
  • translate ambiguous customer problems into production AI architectures
  • rapidly prototype new ideas
  • identify reusable patterns
  • work across the full lifecycle of modern AI systems
  • shipping production systems into enterprise environments
  • measuring business impact
  • continuously improving agents using real-world feedback
  • experience building production AI systems using LLMs
  • experience with modern AI tooling
  • experience designing or evaluating AI systems using quantitative metrics
  • experience building production AI agents or autonomous systems
  • deep understanding of reasoning, retrieval, memory, planning, and tool use
  • experience designing evaluation frameworks for LLMs and agentic system