#data-pipelines (15 Repositories)
Ranked open-source repositories tagged with #data-pipelines, scored by pull request acceptance likelihood and maintainer engagement velocity.
39.1%
56.4h
15 repositories tagged #data-pipelines
feldera/feldera
The Feldera Incremental Computation Engine
wingfoil-io/wingfoil
graph based stream processing framework
conductor-oss/python-sdk
Conductor OSS SDK for Python programming language
bruin-data/bruin
Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
apache/dolphinscheduler
Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code
smart-data-lake/smart-data-lake
Smart Automation Tool for building modern Data Lakes and Data Pipelines
tuva-health/tuva-core
Main repo including core data model, data marts, data quality tests, and terminology sets.
kedro-org/kedro
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
pathwaycom/pathway
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
elementary-data/elementary
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
combust/mleap
MLeap: Deploy ML Pipelines to Production
mage-ai/mage-ai
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Burla-Cloud/burla
The simplest way to scale Python.
dataflint/spark
Drop-in replacement for Apache Spark UI
dagster-io/dagster
An orchestration platform for the development, production, and observation of data assets.