AI Platform Engineer

Supabase Supabase · Data AI · Remote · Growth

Staff AI Platform Engineer to build the execution layer for Supabase's internal AI systems, focusing on an AI-native internal operating system. Responsibilities include building the agent platform, evaluation layer, agents, enforcing governance in code, designing human interaction, owning platform tooling, computing operating measures, and instrumenting the platform's return. The role emphasizes recursive and inversion thinking for robust and trustworthy AI systems.

What you'd actually do

  1. Ship the agent platform to production. An event-triggered queue, a headless model-agnostic runtime (choosing the runtime is an open decision you will close), durable state that survives a failed run, a human review gate, atomic rollback, and complete run logging in the warehouse.
  2. Own the evaluation layer, and switch on the gate that depends on it. Golden suites with behavioral assertions rather than intuition, judge criteria with a written rubric, safety cases that must pass on every run, and a CI gate that blocks a regression from merging. This is the precondition for every agent that does anything beyond read internal data.
  3. Build and register the agent portfolio. Reporting, drafting, linting, triage and question-answering agents across the executive, team-lead and individual-contributor layers, plus a meta layer that observes the platform and improves it.
  4. Enforce governance in code. Risk tiers the pipeline actually enforces, least-privilege credentials per agent, tool-permission gates, a decision audit log, and autonomy classes where the dangerous class has no code path rather than a warning label. Nothing runs without a registered owner, tier, tool grant and human gate.
  5. Design how the system contacts people. A hard interruption budget per person, message bundling instead of a stream of pings, and a structure that gives something useful before it asks for anything. Adoption depends on this more than on any other single design choice.

Skills

Required

  • building agent platforms
  • event-triggered queues
  • model-agnostic runtimes
  • durable state management
  • human review workflows
  • atomic rollback
  • run logging
  • evaluation frameworks
  • golden suites
  • behavioral assertions
  • written rubrics
  • safety case development
  • CI/CD for AI
  • agent development (reporting, drafting, linting, triage, Q&A)
  • meta-agent development
  • code governance
  • risk tiering
  • least-privilege access control
  • tool permissioning
  • audit logging
  • autonomy control
  • system design for human-AI interaction
  • interruption budgeting
  • message bundling
  • platform tooling (compiler, validator, inventory management)
  • data pipelines for operational metrics
  • production data analysis
  • system instrumentation
  • recursive thinking
  • inversion thinking

Nice to have

  • Postgres
  • Supabase platform

What the JD emphasized

  • build the platform that actually runs agents
  • build the evaluation layer that makes any of it trustworthy
  • enforce governance in code
  • risk-tiered
  • human review gate
  • evaluation layer
  • safety cases
  • CI gate
  • least-privilege credentials
  • tool-permission gates
  • decision audit log
  • autonomy classes
  • registered owner
  • tier
  • tool grant
  • human gate
  • hard interruption budget
  • message bundling
  • compute the operating measures from production data
  • instrument the platform's own return
  • recursive thinking
  • inversion thinking
  • agent may never autonomously write a commitment into a shared work system
  • credential in the agent's tool grant physically cannot set an owner, a due date or a status

Other signals

  • building the execution layer for internal AI systems
  • AI-native internal operating system
  • build the platform that actually runs agents
  • build the evaluation layer that makes any of it trustworthy
  • build the agents themselves
  • governance-heavy environment
  • risk-tiered agents
  • enforce governance in code
  • compute the operating measures from production data
  • instrument the platform's own return