#etl (30 Repositories)
Ranked open-source repositories tagged with #etl, scored by pull request acceptance likelihood and maintainer engagement velocity.
54.1%
54.9h
30 repositories tagged #etl
koralium/flowtide
High-performance streaming SQL query engine designed for real-time data processing. Use cases include event-driven architectures, ETL pipelines, and modern data-intensive applications.
apache/hop
Hop Orchestration Platform
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
cocoindex-io/cocoindex
Incremental engine for long horizon agents 🌟 Star if you like it!
ICIJ/extract
A cross-platform command line tool for parallelised content extraction and analysis.
scriptella/scriptella-etl
Scriptella is an open source ETL (Extract-Transform-Load) and script execution tool written in Java.
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
starlake-ai/starlake
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
AbsaOSS/cobrix
A COBOL parser and Mainframe/EBCDIC data source for Apache Spark
apache/devlake
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.
apecloud/ape-dts
ApeCloud's Data Transfer Suite, written in Rust. Provides ultra-fast data replication between MySQL, PostgreSQL, Redis, MongoDB, Kafka and ClickHouse, ideal for disaster recovery (DR) and migration scenarios.
paillave/Etl.Net
Mass processing data with a complete ETL for .net developers
cloudquery/cloudquery
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
PeerDB-io/peerdb
Fast, Simple and a cost effective tool to replicate data from Postgres to Data Warehouses, Queues and Storage
slothflowlabs/duckle
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
nightscape/spark-excel
A Spark plugin for reading and writing Excel files
Open-Source-Legal/OpenContracts
The open document intelligence platform for builders and hackers - DMS for the agentic world
SQLMesh/sqlmesh
Scalable and efficient data transformation framework - backwards compatible with dbt.
airbytehq/airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
apache/flink-cdc
Flink CDC is a streaming data integration tool
TianLangStudio/DataXServer
为DataX(https://github.com/alibaba/DataX) 提供远程多语言调用(ThriftServer,HttpServer) 分布式运行(DataX on YARN) 功能
datavane/tis
Support agile Ontology DataOps Based on Flink, DataX and Flink-CDC with Web-UI
SETL-Framework/setl
A simple Spark-powered ETL framework that just works 🍺
YotpoLtd/metorikku
A simplified, lightweight ETL Framework based on Apache Spark
datamade/data-making-guidelines
:blue_book: Making Data, the DataMade Way
aws-solutions-library-samples/data-lakes-on-aws
Enterprise-grade, production-hardened, serverless data lake on AWS
aelassas/wexflow
Workflow Automation Engine
Swirrl/grafter
Linked Data & RDF Manufacturing Tools in Clojure
stn1slv/awesome-integration
A curated list of awesome system integration software and resources.
dagster-io/dagster
An orchestration platform for the development, production, and observation of data assets.