#data-engineering (30 Repositories)
Ranked open-source repositories tagged with #data-engineering, scored by pull request acceptance likelihood and maintainer engagement velocity.
78.6%
62.9h
30 repositories tagged #data-engineering
bacalhau-project/bacalhau
Community-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.
kanton-bern/hellodata-be
The Open-Source Enterprise Data Platform in a single Portal
rocky-data/rocky
A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.
koralium/flowtide
High-performance streaming SQL query engine designed for real-time data processing. Use cases include event-driven architectures, ETL pipelines, and modern data-intensive applications.
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
datacontract/datacontract-cli
Enforce Data Contracts
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
starlake-ai/starlake
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
DataSQRL/sqrl
Agentic Data Engineering Harness for building data pipelines, data products, data APIs, and data lakes autonomously
DataRecce/recce
The data-validation toolkit for enhanced dbt (data build tool) PR review
dathere/qsv
Blazing-fast Data-Wrangling toolkit
rpbouman/huey
Browser-based data analytics & BI with DuckDB-WASM. Privacy-first: your data stays local; no server required.
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
apache/devlake
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.
bitol-io/open-data-contract-standard
Home of the Open Data Contract Standard (ODCS).
Zipstack/unstract
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
apache/airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
edkreuk/FMD_FRAMEWORK
The Fabric Metadata-Driven Framework (FMD) is a cutting-edge accelerator designed to optimize data handling and utilization.
bitol-io/open-data-product-standard
Home of the Open Data Product Standard (ODPS).
hadziqmtqn/erd-builder-pro
ERD Builder Pro is a database design and documentation tool for developers. Build ERDs, flowcharts, notes, and drawings — all in one workspace
apache/superset
Apache Superset is a Data Visualization and Data Exploration Platform
Zleap-AI/SAG
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
metarank/metarank
A low code Machine Learning personalized ranking service for articles, listings, search results, recommendations that boosts user engagement. A friendly Learn-to-Rank engine
risingwavelabs/risingwave
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
lakehq/sail
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
encero-systems/incan
Incan: a modern, Pythonic language that compiles to Rust! Type-safe, async-friendly, with fixtures, testing, and web/inter-op built in.
slothflowlabs/duckle
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
pyjanitor-devs/pyjanitor
Clean APIs for data cleaning. Python implementation of R package Janitor
Eventual-Inc/Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale