Software Engineer 3

MongoDB MongoDB · Enterprise · VT · Remote · AI Builder Experience

Software Engineer to build and maintain tooling, evaluation systems, and quality gates for agent skills within MongoDB's AI Builder Experience organization. The role focuses on the platform layer around agents, including authoring, distribution, evaluation, monitoring, and improvement of agent skills and behavior.

What you'd actually do

  1. Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them
  2. Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals
  3. Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage
  4. Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression
  5. Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI

Skills

Required

  • 2+ years of experience building production software, developer tools, internal platforms, or automation systems
  • Software engineering fundamentals in API design, testing, error handling, and maintainability
  • Experience building CLIs, libraries, test infrastructure, static analysis, or CI/CD workflows
  • Ability to design systems that are usable by developers and reliable in automation
  • Experience reasoning about correctness and safety with ambiguous input, nondeterministic output, false positives, or untrusted content
  • Comfort in an evolving R&D environment where the right abstraction emerges through prototypes and feedback
  • Written and verbal communication, including explaining technical trade-offs and aligning stakeholders across teams

Nice to have

  • Experience with agentic systems, LLM applications, prompt or rubric-based evaluation, or AI-assisted development
  • Experience building eval harnesses, benchmark datasets, quality metrics, LLM-as-judge workflows, or human-review tooling
  • Experience with Go, Python, JavaScript/TypeScript, Java, or C#
  • Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement
  • Experience moving prototypes into production

What the JD emphasized

  • agent skills
  • agent behavior
  • evaluation
  • quality gates
  • tooling
  • agent-authored content
  • agent skills
  • agent behavior
  • evaluation
  • quality gates
  • tooling
  • agent-authored content

Other signals

  • Agent platform
  • Tooling for agents
  • Evaluation systems for agents
  • Quality gates for agents