#big-data (30 Repositories)
Ranked open-source repositories tagged with #big-data, scored by pull request acceptance likelihood and maintainer engagement velocity.
77.2%
58.3h
30 repositories tagged #big-data
G-Research/ParquetSharp
ParquetSharp is a .NET library for reading and writing Apache Parquet files.
firebolt-db/firebolt-core
Firebolt Core is a free, self-hosted edition of Firebolt's distributed query engine (https://www.firebolt.io/); it provides high-performance data warehousing capabilities that can be deployed anywhere from a single laptop to enterprise datacenters.
apache/ambari
Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.
apache/paimon-rust
Apache Paimon Rust The rust implementation of Apache Paimon.
apache/datafusion-ballista
Apache DataFusion Ballista Distributed Query Engine
vespa-engine/vespa
The AI search platform
StarRocks/starrocks
The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.
apache/paimon
Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.
apache/auron
The Auron accelerator for distributed computing framework (e.g., Spark) leverages native vectorized execution to accelerate query processing
gchq/stroom
Stroom is a highly scalable data storage, processing and analysis platform.
apache/iotdb
Apache IoTDB
apache/zeppelin
Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.
ClickHouse/ClickBench
ClickBench: a Benchmark For Analytical Databases
graphframes/graphframes
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
apache/knox
Mirror of Apache Knox
apache/tsfile
Apache TsFile
apache/datafusion
Apache DataFusion SQL Query Engine
apache/doris-website
Apache Doris Website
ClickHouse/ClickHouse
ClickHouse® is a real-time analytics database management system
apache/ozone
Scalable, reliable, distributed storage system optimized for data analytics and object store workloads.
NVIDIA/cudf-spark
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
kubeflow/mcp-apache-spark-history-server
MCP Server and CLI for Apache Spark History Server. Debug Spark applications from AI agents, scripts, or the terminal.
apache/beam
Apache Beam is a unified programming model for Batch and Streaming data processing.
uxlfoundation/oneDAL
oneAPI Data Analytics Library (oneDAL)
prestodb/presto
The official home of the Presto distributed SQL query engine for big data
TouK/nussknacker
Low-code tool for automating actions on real time data | Stream processing for the users.
apache/calcite
Apache Calcite
trinodb/trino
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
apache/fluss
Apache Fluss is a streaming storage built for real-time analytics.
lakehq/sail
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.