#bigdata (30 Repositories)
Ranked open-source repositories tagged with #bigdata, scored by pull request acceptance likelihood and maintainer engagement velocity.
30.9%
21.8h
30 repositories tagged #bigdata
rustfs/rustfs
🚀2.3x faster than MinIO for 4KB object payloads. RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.
apache/shardingsphere
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
apache/hudi
Upserts, Deletes And Incremental Processing on Big Data.
databendlabs/databend
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
apache/livy
Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.
pbreheny/biglasso
biglasso: Extending Lasso Model Fitting to Big Data in R
dotnet/spark
.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.
gearpump/gearpump
Lightweight real-time big data streaming engine over Akka
volcano-sh/volcano
A Cloud Native Batch System (Project under CNCF)
NationalSecurityAgency/datawave
DataWave is an ingest/query framework that leverages Apache Accumulo to provide fast, secure data access.
jamesmudd/jhdf
A pure Java HDF5 library
apache/avro
Apache Avro is a data serialization system.
AbsaOSS/spline
Data Lineage Tracking And Visualization Solution
apache/celeborn
Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.
byzer-org/byzer-lang
Byzer (former MLSQL): A low-code open-source programming language for data pipeline, analytics and AI.
spotify/big-data-rosetta-code
Code snippets for solving common big data problems in various platforms. Inspired by Rosetta Code
marmotdata/marmot
The open-source context layer for your AI. Catalog your tables, topics, queues and APIs then expose real metadata to your AI agents.
arvados/arvados
An open source platform for managing and analyzing biomedical big data
pingcap/tispark
TiSpark is built for running Apache Spark on top of TiDB/TiKV
achtungsoftware/alarik
High performance distributed S3 compatible object storage focused on speed and designed to be an open alternative to MinIO and RustFS.
MemVerge/splash
Splash, a flexible Spark shuffle manager that supports user-defined storage backends for shuffle data storage and exchange
transferia/transferia
Open Source Cloud Native Ingestion engine
chatnoir-eu/chatnoir-resiliparse
A robust web archive analytics toolkit
aws-samples/aws-etl-orchestrator
A serverless architecture for orchestrating ETL jobs in arbitrarily-complex workflows using AWS Step Functions and AWS Lambda.
shzlw/poli
An easy-to-use BI server built for SQL lovers. Power data analysis in SQL and gain faster business insights.
NewLifeX/AntJob
高吞吐 .NET 分布式任务与实时数据调度平台:时间/数据/消息/Cron/SQL/脚本切片,自动重试与弹性扩缩,回溯补算 + Web 控制台。High‑throughput .NET distributed job & real‑time scheduler with fine‑grained slicing, retries, elastic scaling & web console.
ganweisoft/TOMs
TOMs is a fully open-source, high-performance, systematic, plugin-oriented, and scenario-agnostic general-purpose development framework.
scikit-hep/uproot5
ROOT I/O in pure Python and NumPy.
simbafl/DataWarehouse
从数据仓库到用户画像,从数据建设到数据应用