Senior Software Engineer - Python and Data Ecosystem

ClickHouse ClickHouse · Data AI · Israel +2 · Engineering

Senior Software Engineer to build and maintain Python integrations for ClickHouse, focusing on data ecosystem and AI/LLM tooling like RAG, vector stores, and LLM agents. Role requires hands-on experience as a Data Engineer or Data Scientist with AI/ML in production data contexts.

What you'd actually do

  1. Own and evolve ClickHouse's Python connector and SDK ecosystem, raising the bar on performance, reliability, and API design
  2. Build and maintain integrations with orchestration platforms (Airflow, Dagster, Prefect) and transformation tools (dbt) to enterprise-grade quality standards
  3. Drive the AI/LLM integration strategy: designing connectors and patterns that make ClickHouse a natural fit in RAG architectures, ML feature pipelines, and LLM-powered data applications
  4. Engage actively with the open-source community: triage issues, support contributors, advocate for users, and shape the roadmap based on real-world feedback
  5. Collaborate with Product, Cloud, and other engineering teams to align integration work with broader platform priorities
  6. Bring a practitioner's perspective to roadmap decisions, grounding prioritization in genuine Data Engineer and Data Scientist workflows

Skills

Required

  • Python
  • Data Ecosystem
  • Python connectors
  • SDKs
  • integrations
  • orchestration platforms
  • transformation tools
  • AI/LLM integration strategy
  • RAG architectures
  • ML feature pipelines
  • LLM-powered data applications
  • open-source community engagement
  • Product collaboration
  • Data Engineer workflows
  • Data Scientist workflows
  • Pandas
  • NumPy
  • Pydantic
  • SQL
  • data modeling
  • query optimization
  • OLAP/analytical databases
  • concurrent Python
  • threading
  • multiprocessing
  • async patterns
  • written and verbal communication

Nice to have

  • ClickHouse
  • similar high-performance OLAP platforms
  • JVM ecosystem
  • deploying AI/ML models in production
  • inference APIs
  • vector databases

What the JD emphasized

  • lived the Data Engineer or Data Scientist experience firsthand
  • operated within them
  • hands-on time as a Data Engineer, Data Scientist, or ML Engineer
  • Hands-on experience applying AI/ML in production data-engineering contexts

Other signals

  • building integrations for AI workloads
  • connectors for AI tooling
  • AI-powered data applications
  • vector stores for RAG pipelines
  • backends for LLM-powered agents
  • ML feature stores
  • embedding pipelines
  • retrieval-augmented generation
  • LLM-powered data applications