#data-pipeline (23 Repositories)
Ranked open-source repositories tagged with #data-pipeline, scored by pull request acceptance likelihood and maintainer engagement velocity.
50.3%
48.3h
23 repositories tagged #data-pipeline
rocky-data/rocky
A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
dagucloud/dagu
Self-hostable workflow orchestrator for teams whose main work isn't orchestration. Declarative YAML over your scripts, SSH commands, containers, etc; keep workflows separate from business logic. One binary, no database, runs on limited H/W resources. Alternative to Airflow / Cron / Job Scheduler.
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
starlake-ai/starlake
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
DataSQRL/sqrl
Agentic Data Engineering Harness for building data pipelines, data products, data APIs, and data lakes autonomously
bruin-data/ingestr
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
debezium/debezium
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
rocketride-org/rocketride-server
High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.
datazip-inc/olake
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
clidey/whodb
Where data access meets operational intelligence
slothflowlabs/duckle
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
Hebbian-Robotics/hflow
Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.
snowplow/snowplow
The leader in Customer Data Infrastructure
apache/flink-cdc
Flink CDC is a streaming data integration tool
InfuseAI/piperider
Code review for data in dbt
ModelEngine-Group/DataMate
DataMate is an enterprise-level data processing platform designed for model fine-tuning and RAG retrieval.
Multiwoven/multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
lightfeed/extractor
Use LLMs to robustly extract web data
dataflint/spark
Drop-in replacement for Apache Spark UI
ubisoft/mobydq
:whale: Tool to automate data quality checks on data pipelines
starlake-ai/starflow
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
elementary-data/elementary
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.