#data-integration (18 Repositories)
Ranked open-source repositories tagged with #data-integration, scored by pull request acceptance likelihood and maintainer engagement velocity.
63.1%
148.6h
18 repositories tagged #data-integration
apache/hop
Hop Orchestration Platform
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
starlake-ai/starlake
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
apache/hudi
Upserts, Deletes And Incremental Processing on Big Data.
bruin-data/ingestr
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
apache/devlake
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.
apache/seatunnel
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
apache/airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
cloudquery/cloudquery
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
jitsucom/jitsu
Jitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for modern data teams. Set-up a real-time data pipeline in minutes, not days
clidey/whodb
Where data access meets operational intelligence
slothflowlabs/duckle
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
airbytehq/airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
apache/flink-cdc
Flink CDC is a streaming data integration tool
fluvio-community/fluvio
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
hetio/hetionet
Hetionet: an integrative network of disease
starlake-ai/starflow
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.