Frontier Agents Engineer (applied Ai)

Scale AI Scale AI · Data AI · New York, NY +1 · Enterprise Engineering

This role focuses on designing, evaluating, and deploying production AI agents that leverage frontier models, reasoning techniques, and structured knowledge to solve real-world enterprise problems. The engineer will work across diverse industries and use cases, building multi-agent systems, retrieval pipelines, and ensuring reliability, observability, and safety in production environments. The role requires strong software engineering skills, experience with LLMs and modern AI tooling, and the ability to collaborate directly with customers.

What you'd actually do

  1. Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use.
  2. Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
  3. Design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
  4. Build production-quality AI systems with a strong emphasis on reliability, observability, latency, safety, and cost.
  5. Partner directly with enterprise customers to understand their business, data, and operational challenges.

Skills

Required

  • 4+ years of software engineering, machine learning, or applied AI experience.
  • Strong Python programming skills.
  • Experience building production AI systems using LLMs.
  • Experience with modern AI tooling, including OpenAI, Claude, MCP, agent frameworks, vector databases, or retrieval systems.
  • Strong understanding of machine learning fundamentals and modern language models.
  • Experience designing or evaluating AI systems using quantitative metrics.
  • Excellent communication skills and the ability to work directly with enterprise customers.

Nice to have

  • Experience building production AI agents or autonomous systems.
  • Deep understanding of reasoning, retrieval, memory, planning, and tool use.
  • Experience designing evaluation frameworks for LLMs and agentic systems.

What the JD emphasized

  • production AI agents
  • enterprise customers
  • production AI systems
  • agent architectures
  • evaluation frameworks
  • production environments

Other signals

  • building production AI agents
  • designing and deploying production AI agents
  • architect intelligent systems
  • develop multi-agent systems
  • translate frontier AI research into production systems
  • design rigorous evaluation frameworks
  • build production-quality AI systems
  • design agent guardrails
  • partner directly with enterprise customers
  • build production AI systems using LLMs
  • experience building production AI agents