Senior Software Engineer II - Agentic Intelligence

Honeycomb Honeycomb · Enterprise · United States · Remote · Engineering

Senior Software Engineer II focused on building and shipping production-grade AI agents for an observability platform (Canvas). The role involves end-to-end ownership of agent development, from prototype to production, leveraging a unique high-cardinality data store to enable advanced reasoning capabilities. The engineer will also contribute to the agentic workspace, including memory, spatial awareness, and performance improvements.

What you'd actually do

  1. Design and deliver production-grade agents. Build agents that investigate, reason, and act on live observability data inside Canvas. These agents must be trustworthy to engineers in high pressure situations, including mid-incident. Take one from rough first version to something that holds up under production traffic.
  2. Own the agent work; support the whole product. Scope, build, ship, and maintain the agents including the evals that tell you whether they got better or are just different. This role is agent-focused and also includes some fullstack development.
  3. Build agents only Honeycomb can build. Use a data store that returns high-cardinality queries in seconds to reason over signal a conventional backend can't serve at this fidelity correlating across services, drilling into a single trace, comparing before and after a deploy.
  4. Extend the surface, and decide what's next. Ship new capability into Canvas, the MCP server, and Canvas Skills memory, spatial awareness, a faster Bedrock loop and make the case for what comes after with working code. Distinguish hype from signal in a field with plenty of both.
  5. Define what "good" means for agents here. Set the bar: measurable against real evals, maintainable, and honest about their limits.

Skills

Required

  • AI and agent engineering experience
  • shipped LLM-based systems in production
  • agent systems design
  • end-to-end ownership
  • agent architecture depth
  • product judgment

Nice to have

  • Observability or developer-tools background
  • familiarity with eval frameworks
  • agent tooling
  • RAG
  • prompt engineering

What the JD emphasized

  • production-grade agents
  • production traffic
  • production
  • evals
  • agent architecture depth
  • product judgment

Other signals

  • AI agents
  • LLM-based systems
  • agentic workspace
  • production-grade agents
  • evals