#hadoop (30 Repositories)
Ranked open-source repositories tagged with #hadoop, scored by pull request acceptance likelihood and maintainer engagement velocity.
35.3%
42.4h
30 repositories tagged #hadoop
apache/ozone
Scalable, reliable, distributed storage system optimized for data analytics and object store workloads.
Tencent/APIJSON
🏆 Real-Time no-code, powerful and secure ORM 🚀 providing APIs and Docs without coding by Backend, and Frontend(Client) can customize response JSONs 🏆 实时 零代码、全功能、强安全 ORM 库 🚀 后端接口和文档零代码,前端(客户端) 定制返回 JSON 的数据和结构
linkedin/venice
Venice, Derived Data Platform for Planet-Scale Workloads.
prestodb/presto
The official home of the Presto distributed SQL query engine for big data
trinodb/trino
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
apache/carbondata
High performance data store solution
apache/calcite
Apache Calcite
smart-data-lake/smart-data-lake
Smart Automation Tool for building modern Data Lakes and Data Pipelines
apache/drill
Apache Drill is a distributed MPP query layer for self describing data
wgzhao/Addax
A fast and versatile ETL tool that can transfer data between RDBMS and NoSQL seamlessly
apache/wayang
Apache Wayang is the first cross-platform data processing system.
AbsaOSS/spline
Data Lineage Tracking And Visualization Solution
apache/hadoop
Apache Hadoop
GridProtectionAlliance/openPDC
Open Source Phasor Data Concentrator
apache/doris-website
Apache Doris Website
apache/kyuubi
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
aliyun/aliyun-emapreduce-datasources
Extended datasource support for Spark/Hadoop on Aliyun E-MapReduce.
HariSekhon/DevOps-Bash-tools
1200+ DevOps Bash Scripts - AWS, GCP, Kubernetes, Docker, CI/CD, APIs, SQL, PostgreSQL, MySQL, Hive, Impala, Kafka, Hadoop, Jenkins, GitHub, GitLab, BitBucket, Azure DevOps, TeamCity, Spotify, MP3, LDAP, Code/Build Linting, pkg mgmt for Linux, Mac, Python, Perl, Ruby, NodeJS, Golang, Advanced dotfiles: .bashrc, .vimrc, .gitconfig, .screenrc, tmux..
mjakubowski84/parquet4s
Read and write Parquet in Scala. Use Scala classes as schema. No need to start a cluster.
hortonworks/cloudbreak
CDP Public Cloud is an integrated analytics and data management platform deployed on cloud services. It offers broad data analytics and artificial intelligence functionality along with secure user access and data governance features.
deeplearning4j/deeplearning4j
Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular and tiny c++ library for running math code and a java based math library on top of the core c++ library. Also includes samediff: a pytorch/tensorflow like library for running deep learn...
HariSekhon/HAProxy-configs
80+ HAProxy Configs for Hadoop, Big Data, NoSQL, Docker, Kubernetes, Elasticsearch, SolrCloud, HBase, MySQL, PostgreSQL, Apache Drill, Hive, Presto, Impala, Hue, ZooKeeper, SSH, RabbitMQ, Redis, Riak, Cloudera, OpenTSDB, InfluxDB, Prometheus, Kibana, Graphite, Rancher etc.
elasticluster/elasticluster
Create clusters of VMs on the cloud and configure them with Ansible.
fancyChuan/bigdata-hub
数据建设与大数据技术知识体系,包含hadoop、hive、spark、flink主流框架和系列框架,数据中台、数据湖、数据治理、数仓建设、数据化转型等
spotify/luigi
Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.
Netflix/iceberg
Iceberg is a table format for large, slow-moving tabular data
kanyun-inc/ytk-learn
Ytk-learn is a distributed machine learning library which implements most of popular machine learning algorithms(GBDT, GBRT, Mixture Logistic Regression, Gradient Boosting Soft Tree, Factorization Machines, Field-aware Factorization Machines, Logistic Regression, Softmax).
whoiszxl/shopzz
后端使用 SpringCloud Alibaba 开发,移动端使用 React Native 构建,管理后台使用 Arco Design 进行构建,并在支付上接入数字货币(比特币、以太坊UDST、平台Token)支付,后端采用 Hadoop 与 Flink 等大数据框架构建实时计算与离线计算体系。