Python interface to Hive and Presto.
PythonOthersteady
8 projectsBatch Processing
Python interface to Hive and Presto.
PythonOthersteady
A cloud native data pipeline and transformation toolkit written in Go.
GoMIT Licensesteady
A free & cross platform monitoring tool (Spark UI / Spark History Server alternative).
ScalaOtherslowingarchived
Scalable machine learning library for Hive/Hadoop.
JavaApache License 2.0dormantarchived
Connecting Apache Spark with different data stores. Deprecated.
JavaApache License 2.0dormantarchived
Personal genome analysis toolkit with Python scripts analyzing raw DNA data across 17 categories (health risks, ancestry, pharmacogenomics, nutrition, psychology, etc.) and generating a terminal-style single-page HTML visualization.
PythonMIT Licensesteady
Pure-Go classic machine learning toolkit and data engineering utilities. Eight algorithms with zero external dependencies.
GoMIT Licensesteady
A light-weight engine for general-purpose data processing including both batch and stream analytics. It is based on a novel unique data model, which represents data via functions and processes data via columns operations as opposed to having only set operations in conventional approaches like MapReduce or SQL.
JavaMIT Licensedormant
Back to Awesome DataEngineeringSearch Awesome DataEngineering →