Senior Software Engineer, Assistant Engineering

Airbnb Airbnb · Consumer · United States · Software Engineering

Senior Software Engineer on the AI Assistant Product Evaluation team responsible for designing and building scalable data systems and evaluation platforms for Airbnb's agentic AI products. This role involves working at the intersection of backend engineering, data systems, AI evaluation, and product quality to measure, monitor, and improve the quality of AI products.

What you'd actually do

  1. Design and productionize scalable data systems that support AI evaluation, metric computation, observability, and feedback loops for agentic AI products.
  2. Build data models, schemas, and processing pipelines for agentic AI interactions, supporting reliable logging, retrieval, metric computation, and long-term evaluation dataset management.
  3. Work closely with Core Modeling engineers to understand pain points in the LLM evaluation process, and develop LLM-as-a-judge solutions and data pipelines to address metric-related challenges in a scalable and efficient way.
  4. Collaborate with machine learning infrastructure engineering teams to evolve how we build and test evaluation framework for Airbnb Conversational AI products.
  5. Lead all phases of software development including architecture design, implementation and testing.

Skills

Required

  • 5+ years of industry experience as a software engineer, backend engineer, platform engineer, or data-focused software engineer building production systems.
  • BS, MS, or PhD in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • Strong programming skills in Python, with experience building production-quality software, libraries, frameworks, or data processing systems.
  • Experience designing and operating scalable data pipelines using Airflow or similar orchestration frameworks; experience with Spark, Flink, Kafka, Trino/Presto, Hive, Iceberg, or similar big data technologies is a strong plus.
  • Strong understanding of data modeling, schema design, data quality, partitioning, indexing, storage formats, and tradeoffs for large-scale analytical and operational data systems.
  • Experience building data layers, evaluation systems, feedback loops, experimentation platforms, observability tools, or AI/ML platform capabilities.
  • Ability to analyze complex datasets, identify data quality issues, debug inconsistencies, and translate findings into actionable engineering or product decisions.
  • Strong system design skills, including experience building reliable, extensible, maintainable systems with clear APIs, testing strategies, and operational ownership.
  • Solid understanding of data structures and algorithms, with the ability to make practical engineering tradeoffs for performance, scalability, and maintainability.
  • Familiarity with AI/ML system concepts such as model evaluation, offline evaluation, online monitoring, model quality metrics, human-in-the-loop workflows, experimentation, or model deployment.
  • Proven ability to work cross-functionally with modeling engineers, product managers, data scientists, infrastructure teams, and operations partners to deliver end-to-end solutions.
  • Excellent communication with the ability to drive alignment, set technical direction, and raise engineering quality across a team.

Nice to have

  • Spark, Flink, Kafka, Trino/Presto, Hive, Iceberg, or similar big data technologies

What the JD emphasized

  • scalable data systems
  • evaluation platforms
  • agentic AI products
  • AI evaluation
  • product quality
  • AI quality measurable
  • debuggable
  • actionable
  • LLM evaluation process
  • LLM-as-a-judge solutions
  • evaluation framework

Other signals

  • AI Assistant
  • agentic AI products
  • evaluation systems
  • data foundations
  • observability tools
  • AI product quality
  • model and agent behavior
  • identify regressions
  • feedback loop
  • product iteration
  • scalable data systems
  • evaluation platforms
  • backend engineering
  • data systems
  • AI evaluation
  • product quality
  • data-rich systems
  • product interactions
  • model outputs
  • human feedback
  • evaluation results
  • reliable signals
  • product and model development
  • AI quality measurable
  • debuggable
  • actionable
  • development lifecycle
  • modeling teams
  • product managers
  • data scientists
  • operations teams
  • ambiguous AI evaluation needs
  • robust systems
  • trustworthy AI experiences
  • productionize scalable data systems
  • AI evaluation
  • metric computation
  • observability
  • feedback loops
  • agentic AI products
  • data models
  • schemas
  • processing pipelines
  • agentic AI interactions
  • reliable logging
  • retrieval
  • metric computation
  • long-term evaluation dataset management
  • Core Modeling engineers
  • LLM evaluation process
  • LLM-as-a-judge solutions
  • data pipelines
  • metric-related challenges
  • scalable and efficient way
  • machine learning infrastructure engineering teams
  • evaluation framework
  • Airbnb Conversational AI products
  • architecture design
  • implementation and testing
  • cross-functional partners
  • product managers
  • operations
  • data scientists
  • business impact
  • prioritize requirements
  • machine learning systems
  • data pipelines
  • engineering decisions
  • quantify impact
  • engineering excellence
  • high-quality code
  • operational reliability
  • sharing knowledge