Back to Topics Directory
Topic Hub

#bigdata (30 Repositories)

Ranked open-source repositories tagged with #bigdata, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

30.9%

Avg Review Latency

21.8h

Filter by language

30 repositories tagged #bigdata

S TierRust 31.5k 1 GFIs

rustfs/rustfs

🚀2.3x faster than MinIO for 4KB object payloads. RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.

93.7%
Merge Rate
2d
First Review
79%
1st-Timers
51
Maintainers
S TierJava 20.8k

apache/shardingsphere

Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.

90.7%
Merge Rate
2d
First Review
69%
1st-Timers
25
Maintainers
A TierJava 6.2k 19 GFIs

apache/hudi

Upserts, Deletes And Incremental Processing on Big Data.

65.7%
Merge Rate
14h
First Review
61%
1st-Timers
41
Maintainers
A TierRust 9.4k 1 GFIs

databendlabs/databend

Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.

79.0%
Merge Rate
5d
First Review
79%
1st-Timers
19
Maintainers
A TierScala 959

apache/livy

Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.

60.7%
Merge Rate
3h
First Review
75%
1st-Timers
9
Maintainers
B TierC++ 115

pbreheny/biglasso

biglasso: Extending Lasso Model Fitting to Big Data in R

100.0%
Merge Rate
16h
First Review
100%
1st-Timers
0
Maintainers
A TierC# 2.1k

dotnet/spark

.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.

100.0%
Merge Rate
<1h
First Review
100%
1st-Timers
2
Maintainers
B TierScala 755

gearpump/gearpump

Lightweight real-time big data streaming engine over Akka

49.5%
Merge Rate
1d
First Review
50%
1st-Timers
1
Maintainers
B TierGo 5.9k 11 GFIs

volcano-sh/volcano

A Cloud Native Batch System (Project under CNCF)

44.5%
Merge Rate
3d
First Review
40%
1st-Timers
53
Maintainers
B TierJava 721

NationalSecurityAgency/datawave

DataWave is an ingest/query framework that leverages Apache Accumulo to provide fast, secure data access.

68.3%
Merge Rate
3d
First Review
55%
1st-Timers
15
Maintainers
B TierJava 172

jamesmudd/jhdf

A pure Java HDF5 library

76.2%
Merge Rate
3d
First Review
100%
1st-Timers
3
Maintainers
B TierJava 3.3k

apache/avro

Apache Avro is a data serialization system.

47.4%
Merge Rate
2h
First Review
33%
1st-Timers
32
Maintainers
B TierScala 667

AbsaOSS/spline

Data Lineage Tracking And Visualization Solution

50.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
B TierJava 1.1k

apache/celeborn

Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.

1.1%
Merge Rate
24h
First Review
0%
1st-Timers
28
Maintainers
D TierScala 1.8k

byzer-org/byzer-lang

Byzer (former MLSQL): A low-code open-source programming language for data pipeline, analytics and AI.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierScala 297

spotify/big-data-rosetta-code

Code snippets for solving common big data problems in various platforms. Inspired by Rosetta Code

0.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
D TierGo 607

marmotdata/marmot

The open-source context layer for your AI. Catalog your tables, topics, queues and APIs then expose real metadata to your AI agents.

0.0%
Merge Rate
5d
First Review
0%
1st-Timers
3
Maintainers
D TierGo 420

arvados/arvados

An open source platform for managing and analyzing biomedical big data

0.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
D TierScala 889

pingcap/tispark

TiSpark is built for running Apache Spark on top of TiDB/TiKV

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierSwift 515

achtungsoftware/alarik

High performance distributed S3 compatible object storage focused on speed and designed to be an open alternative to MinIO and RustFS.

0.0%
Merge Rate
2d
First Review
0%
1st-Timers
1
Maintainers
D TierScala 131

MemVerge/splash

Splash, a flexible Spark shuffle manager that supports user-defined storage backends for shuffle data storage and exchange

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 193

transferia/transferia

Open Source Cloud Native Ingestion engine

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 144

chatnoir-eu/chatnoir-resiliparse

A robust web archive analytics toolkit

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierScala 402

pierre94/flink-notes

flink学习笔记

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 345

aws-samples/aws-etl-orchestrator

A serverless architecture for orchestrating ETL jobs in arbitrarily-complex workflows using AWS Step Functions and AWS Lambda.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJava 2.0k

shzlw/poli

An easy-to-use BI server built for SQL lovers. Power data analysis in SQL and gain faster business insights.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC# 427

NewLifeX/AntJob

高吞吐 .NET 分布式任务与实时数据调度平台:时间/数据/消息/Cron/SQL/脚本切片,自动重试与弹性扩缩,回溯补算 + Web 控制台。High‑throughput .NET distributed job & real‑time scheduler with fine‑grained slicing, retries, elastic scaling & web console.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierCSS 441

ganweisoft/TOMs

TOMs is a fully open-source, high-performance, systematic, plugin-oriented, and scenario-agnostic general-purpose development framework.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 270

scikit-hep/uproot5

ROOT I/O in pure Python and NumPy.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 648

simbafl/DataWarehouse

从数据仓库到用户画像,从数据建设到数据应用

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers