Software Engineer 3

MongoDB MongoDB · Enterprise · VT · Remote · AI Builder Experience

Software Engineer to build and maintain tooling, evaluation systems, and quality gates for MongoDB's agent skills. The role focuses on the platform layer around agents, including authoring, distribution, evaluation, monitoring, and improvement of agent skills and behavior. This involves creating CLIs, libraries, evaluation harnesses, and CI workflows to ensure the quality and safety of AI-generated content.

What you'd actually do

  1. Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them
  2. Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals
  3. Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage
  4. Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression
  5. Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI

Skills

Required

  • 2+ years of experience building production software, developer tools, internal platforms, or automation systems
  • Software engineering fundamentals in API design, testing, error handling, and maintainability
  • Experience building CLIs, libraries, test infrastructure, static analysis, or CI/CD workflows
  • Ability to design systems that are usable by developers and reliable in automation
  • Experience reasoning about correctness and safety with ambiguous input, nondeterministic output, false positives, or untrusted content
  • Comfort in an evolving R&D environment where the right abstraction emerges through prototypes and feedback
  • Written and verbal communication, including explaining technical trade-offs and aligning stakeholders across teams

Nice to have

  • Experience with agentic systems, LLM applications, prompt or rubric-based evaluation, or AI-assisted development
  • Experience building eval harnesses, benchmark datasets, quality metrics, LLM-as-judge workflows, or human-review tooling
  • Experience with Go, Python, JavaScript/TypeScript, Java, or C#
  • Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement
  • Experience moving prototypes into production

What the JD emphasized

  • build and maintain
  • design evaluation datasets and workflows
  • build agent metrics and observability
  • design safety and quality gates
  • build code-generation quality checks
  • investigate real failures

Other signals

  • building tooling for agent skills
  • designing evaluation datasets and workflows
  • building agent metrics and observability
  • designing safety and quality gates for agent-authored content
  • integrating tooling into CI workflows