#spark (30 Repositories)
Ranked open-source repositories tagged with #spark, scored by pull request acceptance likelihood and maintainer engagement velocity.
78.6%
72.6h
30 repositories tagged #spark
ascii-supply-networks/dagster-slurm
Dagster SLURM integration
nielsbasjes/splittablegzip
Splittable Gzip codec for Hadoop
neo4j/neo4j-spark-connector
Neo4j Connector for Apache Spark, which provides bi-directional read/write access to Neo4j from Spark, using the Spark DataSource APIs
MigoXLab/dingo
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
kamu-data/kamu-cli
Next-generation decentralized data lakehouse and a multi-party stream processing network
apache/paimon
Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.
apache/auron
The Auron accelerator for distributed computing framework (e.g., Spark) leverages native vectorized execution to accelerate query processing
apache/zeppelin
Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.
starlake-ai/starlake
Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.
AbsaOSS/cobrix
A COBOL parser and Mainframe/EBCDIC data source for Apache Spark
graphframes/graphframes
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
microsoft/SynapseML
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
alexarchambault/ammonite-spark
Run spark calculations from Ammonite
lakesoul-io/LakeSoul
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.
projectnessie/nessie
Nessie: Transactional Catalog for Data Lakes with Git-like semantics
apache/datafusion-comet
Apache DataFusion Comet Spark Accelerator
JohnSnowLabs/spark-nlp
State of the Art Natural Language Processing
almond-sh/almond
A Scala kernel for Jupyter
apache/livy
Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.
NVIDIA/cudf-spark
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
dotnet/spark
.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.
awslabs/data-on-eks
DoEKS is a tool to build, deploy and scale Data Platforms on Amazon EKS
lakehq/sail
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
apache/carbondata
High performance data store solution
broadinstitute/gatk
Official code repository for GATK versions 4 and up
delta-io/delta-sharing
An open protocol for secure data sharing
zinggAI/zingg
Scalable master data management, identity resolution, entity resolution, and deduplication using ML
awslabs/deequ
Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.
apache/amoro
Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.
igniterealtime/Spark
Cross-platform real-time collaboration client optimized for business and organizations.