Back to Topics Directory
Topic Hub

#data-pipeline (23 Repositories)

Ranked open-source repositories tagged with #data-pipeline, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

50.3%

Avg Review Latency

48.3h

Filter by language

23 repositories tagged #data-pipeline

S TierRust 294 1 GFIs

rocky-data/rocky

A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.

97.1%
Merge Rate
4d
First Review
83%
1st-Timers
2
Maintainers
A TierGo 608

ConduitIO/conduit

Conduit streams data between data stores. Kafka Connect replacement. No JVM required.

80.3%
Merge Rate
2h
First Review
50%
1st-Timers
2
Maintainers
A TierGo 3.8k

dagucloud/dagu

Self-hostable workflow orchestrator for teams whose main work isn't orchestration. Declarative YAML over your scripts, SSH commands, containers, etc; keep workflows separate from business logic. One binary, no database, runs on limited H/W resources. Alternative to Airflow / Cron / Job Scheduler.

94.2%
Merge Rate
4d
First Review
77%
1st-Timers
12
Maintainers
A TierRust 962

estuary/flow

🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊

74.9%
Merge Rate
15h
First Review
78%
1st-Timers
20
Maintainers
B TierScala 210

starlake-ai/starlake

Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.

94.1%
Merge Rate
4d
First Review
100%
1st-Timers
0
Maintainers
A TierJava 222

DataSQRL/sqrl

Agentic Data Engineering Harness for building data pipelines, data products, data APIs, and data lakes autonomously

85.1%
Merge Rate
2d
First Review
67%
1st-Timers
5
Maintainers
A TierGo 3.9k

bruin-data/ingestr

ingestr is a CLI tool to copy data between any databases with a single command seamlessly.

86.3%
Merge Rate
6d
First Review
83%
1st-Timers
13
Maintainers
A TierJava 13.0k

debezium/debezium

Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.

79.3%
Merge Rate
3d
First Review
68%
1st-Timers
42
Maintainers
B TierPython 7.0k 3 GFIs

rocketride-org/rocketride-server

High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.

64.2%
Merge Rate
2d
First Review
36%
1st-Timers
31
Maintainers
B TierGo 1.4k 6 GFIs

datazip-inc/olake

OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.

62.0%
Merge Rate
2d
First Review
44%
1st-Timers
27
Maintainers
B TierGo 5.0k 1 GFIs

clidey/whodb

Where data access meets operational intelligence

85.6%
Merge Rate
13d
First Review
100%
1st-Timers
2
Maintainers
B TierRust 1.2k

slothflowlabs/duckle

Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.

86.7%
Merge Rate
3d
First Review
100%
1st-Timers
1
Maintainers
B TierPython 147 7 GFIs

Hebbian-Robotics/hflow

Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.

16.7%
Merge Rate
7h
First Review
0%
1st-Timers
8
Maintainers
B TierScala 7.0k

snowplow/snowplow

The leader in Customer Data Infrastructure

100.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
B TierJava 6.5k

apache/flink-cdc

Flink CDC is a streaming data integration tool

50.2%
Merge Rate
4d
First Review
52%
1st-Timers
30
Maintainers
D TierPython 495

InfuseAI/piperider

Code review for data in dbt

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 367

ModelEngine-Group/DataMate

DataMate is an enterprise-level data processing platform designed for model fine-tuning and RAG retrieval.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRuby 1.7k

Multiwoven/multiwoven

🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 319

lightfeed/extractor

Use LLMs to robustly extract web data

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 490

dataflint/spark

Drop-in replacement for Apache Spark UI

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierVUVue 257

ubisoft/mobydq

:whale: Tool to automate data quality checks on data pipelines

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierScala 212

starlake-ai/starflow

Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierHTML 2.4k

elementary-data/elementary

The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Data-pipeline Open Source Repositories & C-Rank™ | GetMerged