#parquet (30 Repositories)
Ranked open-source repositories tagged with #parquet, scored by pull request acceptance likelihood and maintainer engagement velocity.
50.1%
118.0h
30 repositories tagged #parquet
apecloud/myduckserver
Unified MySQL, Postgres & FlightSQL Server, Powered by DuckDB.
G-Research/ParquetSharp
ParquetSharp is a .NET library for reading and writing Apache Parquet files.
infino-ai/infino
Fast search engine on object storage, with full text search, vectors, and SQL, natively on Parquet.
Basekick-Labs/arc
Open, SQL-native time-series database for telemetry you need to keep. 34M+ records/sec ingestion, 8M+ rows/sec queries. InfluxDB Line Protocol and Telegraf compatible. Open Parquet on your storage. Single binary. S3/Azure native. Air-gap ready. AGPL-3.0.
firebolt-db/firebolt-core
Firebolt Core is a free, self-hosted edition of Firebolt's distributed query engine (https://www.firebolt.io/); it provides high-performance data warehousing capabilities that can be deployed anywhere from a single laptop to enterprise datacenters.
nao1215/filesql
loads CSV, TSV, LTSV, JSON, JSONL, Parquet, XLSX, ACH, and Fedwire files into SQLite; includes prep and frame for cleanup and in-memory transforms
parquet-go/parquet-go
High-performance Go package to read and write Parquet files
rilldata/rill
The fastest business intelligence tool for humans and agents.
hardwood-hq/hardwood
A fast minimal dependency implementation of Apache Parquet
hangxie/parquet-tools
A utility to deal with Parquet data
hyparam/hyparquet
parquet file parser for javascript
datazip-inc/olake
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
Hebbian-Robotics/hflow
Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.
apache/arrow-rs
Official Rust implementation of Apache Arrow
apache/arrow
Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics
apache/drill
Apache Drill is a distributed MPP query layer for self describing data
nisshi-io/nisshi
Apache Kafka® compatible broker with S3, PostgreSQL, SQLite, Apache Iceberg and Delta Lake
apache/parquet-java
Apache Parquet Java
julien040/anyquery
One SQL interface for 60+ tools (e.g., GitHub, Notion, Airtable). Plug into any LLM through MCP.
spotify/ratatool
A tool for data sampling, data generation, and data diffing
JuliaIO/Parquet.jl
Julia implementation of Parquet columnar file format reader
mjakubowski84/parquet4s
Read and write Parquet in Scala. Use Scala classes as schema. No need to start a cluster.
viggy28/streambed
Stream Postgres to Apache Iceberg on S3 via logical replication, queryable over the Postgres wire protocol.
manojkarthick/pqrs
Command line tool for inspecting Parquet files
caioricciuti/duck-ui
The fully open-source DuckDB workbench that runs in your browser. SQL editor, notebooks, charts, AI assistant. No install, no signup, no backend — your data never leaves the tab.
Netflix/iceberg
Iceberg is a table format for large, slow-moving tabular data
XiangpengHao/liquid-cache
Pushdown cache for DataFusion
hyparam/icebird
Icebird: JavaScript Iceberg Client
bigdatagenomics/adam
ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.
jorgecarleitao/parquet2
Fastest and safest Rust implementation of parquet. `unsafe` free. Integration-tested against pyarrow