Senior Data Platform Engineer

Pinecone Pinecone · Data AI · New York, NY · R&D

Senior Data Engineer to own and grow the systems that power how Pinecone understands its business. This role will design and operate the ingest, transform, orchestration, and metrics layers that feed analysts, executives, and the Board. Responsibilities include building ingestion and transform layers, operating the orchestration platform, establishing the business-context and metrics layer, managing infrastructure cost and performance, leading company-level analyses, enabling self-serve for other teams, and setting standards for AI-assisted data workflows.

What you'd actually do

  1. Own and build the ingestion layer. **Design, deploy, and scale pipelines that pull from third-party APIs, internal services, and SaaS tools into BigQuery. Add new sources as the business demands.
  2. Own and build the transform layer. **Develop and maintain our DBT project, including staging, intermediate, and marts. Maintain core business datasets: users, organizations, indexes, accounts, usage, revenue. Write tests, snapshots, and documentation. Drive data quality and trust.
  3. Own and build the orchestration platform. **Operate the Airflow-on-Kubernetes environment that runs our ingest and DBT workloads. Improve reliability, scalability, observability, and CI/CD.
  4. Establish and maintain the business-context and metrics layer. **Curate metric definitions and documentation that feed both human analysts and agents.
  5. Set the standard for AI-assisted data workflow. **Establish best AI practices and patterns that enable a small data team to operate with outsized leverage.

Skills

Required

  • 4+ years building and operating data pipelines in production.
  • Strong SQL, with comfort in BigQuery (or Snowflake/Redshift) writing non-trivial analytical queries, optimizing performance, and reasoning about correctness.
  • Strong coding skills, with comfort writing ETL/rETL, consuming services and integrations against REST/GraphQL APIs, and producing clean code that others can reuse and maintain.
  • Experience with a modern orchestrator (Airflow, Dagster, Prefect, or similar) running containerized workloads.
  • Comfort with Docker, Kubernetes, and modern cloud infrastructure best practices.
  • Experience integrating systems, pulling data between APIs, databases, and warehouses; handling auth, pagination, schema drift, and incremental loads.
  • Hands-on experience using AI coding tools (Claude Code, Cursor, or similar) as part of your workflow.
  • Ability to design, build, and own systems end-to-end in a highly autonomous environment.

Nice to have

  • Production DBT experience: layered models, tests, snapshots, macros, deferred builds.
  • Experience working with a semantic layer, metrics layer (DBT Semantic Layer, Cube, LookML).
  • Comfortable with exploratory analysis, designing experiments and A/B tests, basic statistical modeling, and separating signal from noise in messy data.
  • Exposure to building AI agents or applications.
  • Infrastructure-as-code (Terraform, Pulumi, or similar).

What the JD emphasized

  • AI-assisted data workflow