Remedy small files by combining them into larger ones.
☆196Jul 1, 2022Updated 4 years ago
Alternatives and similar repositories for filecrush
Users that are interested in filecrush are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hadoop utility to compact small files☆18Feb 16, 2026Updated 6 months ago
- Memory / Configuration Calculator for Hive LLAP☆14Jul 18, 2020Updated 6 years ago
- Remedy small files by combining them into larger ones.☆23Oct 31, 2018Updated 7 years ago
- Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.☆1,133Apr 10, 2023Updated 3 years ago
- A service which allows Hive Metastore Listeners to be deployed outside of the Hive Metastore Service☆13Jun 30, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆34Jan 13, 2019Updated 7 years ago
- LinkedIn's previous generation Kafka to HDFS pipeline.☆878Aug 27, 2020Updated 6 years ago
- Tool for gathering blocks and replicas meta data from HDFS. It also builds a heat map showing how replicas are distributed along disks an…☆55May 9, 2017Updated 9 years ago
- Visualize your HDFS cluster usage☆228Oct 13, 2020Updated 5 years ago
- Hadoop library for large-scale data processing, now an Apache Incubator project☆581Jul 8, 2014Updated 12 years ago
- Utility to easily copy files into HDFS☆70Feb 11, 2020Updated 6 years ago
- SQL Windowing Functions for Hadoop☆65Jun 20, 2022Updated 4 years ago
- A set of Hadoop utilities to make working with Hadoop a little easier.☆26Feb 11, 2020Updated 6 years ago
- Tools for Hadoop☆24Feb 27, 2012Updated 14 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A library for strong, schema based conversion between 'natural' JSON documents and Avro☆18Mar 5, 2024Updated 2 years ago
- Approximate cardinality estimation with HyperLogLog, as a Hive function☆42Dec 17, 2012Updated 13 years ago
- Pig on Apache Spark☆82Mar 23, 2015Updated 11 years ago
- A small project to show how to add lineage to Atlas when using Spark as ETL tool☆12Nov 29, 2016Updated 9 years ago
- Vagrant files creating multi-node virtual Hadoop clusters with or without security.☆67May 13, 2020Updated 6 years ago
- INACTIVE: A daemon to transfer syslog messages to Apache Kafka.☆24Mar 30, 2017Updated 9 years ago
- Tool which generates Avro schemas and Java bindings from XML schemas.☆40Aug 3, 2020Updated 6 years ago
- JUnit integration for testing the Apache Hive Metastore and HiveServer2 Thrift APIs☆26Jul 22, 2025Updated last year
- ☆29Nov 17, 2014Updated 11 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Mahout vector encoding for pig☆53Nov 20, 2022Updated 3 years ago
- Hive + Avro. Serde for working with Avro in Hive☆60Dec 16, 2023Updated 2 years ago
- Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming☆26Sep 10, 2013Updated 12 years ago
- Framework that makes processing arbitrary binary data in Hadoop easier☆22Apr 8, 2013Updated 13 years ago
- GeoIP Functions for hive☆49Oct 13, 2020Updated 5 years ago
- Hannibal is tool to help monitor and maintain HBase-Clusters that are configured for manual splitting.☆172Dec 22, 2017Updated 8 years ago
- Pimped Fork of DataStax Brisk Distribution☆17Oct 31, 2011Updated 14 years ago
- Example MapReduce jobs in Java, Hive, Pig, and Hadoop Streaming that work on Avro data.☆115Nov 12, 2015Updated 10 years ago
- Unit test framework for hive and hive-service☆65Jun 29, 2022Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- An app built on Cloudera Enterprise for tracking metrics of jobs that run in YARN framework☆13Feb 5, 2016Updated 10 years ago
- hRaven collects run time data and statistics from MapReduce jobs in an easily queryable format☆129Jan 14, 2022Updated 4 years ago
- Hive I/O Library☆67Oct 28, 2021Updated 4 years ago
- Sample Python code for working with the HBase REST interface☆24Jul 25, 2013Updated 13 years ago
- Examples of running hadoop clusters with pallet☆24Apr 18, 2014Updated 12 years ago
- 项目中保留了向开源社区提交过的patch☆16Oct 22, 2017Updated 8 years ago
- Jumbune, an open source BigData APM & Data Quality Management Platform for Data Clouds. Enterprise feature offering is available at http:…☆73Jan 1, 2023Updated 3 years ago