An example of using Avro and Parquet in Spark SQL
☆60Nov 16, 2015Updated 10 years ago
Alternatives and similar repositories for avro-parquet-spark-example
Users that are interested in avro-parquet-spark-example are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Example project to show how to use Spark to read and write Avro/Parquet files☆50Aug 21, 2013Updated 13 years ago
- Scripts used to setup a Spark cluster on EC2☆21Mar 24, 2016Updated 10 years ago
- Avro Data Source for Apache Spark☆536Dec 19, 2018Updated 7 years ago
- Spark GCE Script Helps you deploy Spark cluster on Google Cloud.☆43May 30, 2015Updated 11 years ago
- Spark Extension : ML transformers, SQL aggregations, etc that are missing in Apache Spark☆144Jan 26, 2016Updated 10 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Scriptable scheduler for periodical Hadoop workflows☆22Feb 1, 2018Updated 8 years ago
- Embedded Kafka for testing and quick prototyping.☆14Apr 19, 2016Updated 10 years ago
- HBase DSL for Scala with MapReduce support☆128Jan 4, 2018Updated 8 years ago
- Sample App. Amazon Product Descriptions Wordcloud. Spark Streaming, Algebird, Storehaus, Redis, Scala Scraper, OpenNLP, Play Framework, D…☆12Nov 9, 2015Updated 10 years ago
- ☆10Aug 28, 2018Updated 8 years ago
- Simple Spark app that reads and writes Avro data☆31Apr 13, 2015Updated 11 years ago
- 分类模型☆15Apr 19, 2018Updated 8 years ago
- ☆14Aug 23, 2015Updated 11 years ago
- Example of reading/writing Excel files from Pandas/Python☆14Dec 10, 2014Updated 11 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Efficient, distributed downloads of large files from S3 to HDFS using Spark.☆17Apr 26, 2017Updated 9 years ago
- Hadoop for archiving email☆22Sep 29, 2011Updated 14 years ago
- Choosing a fantasy football team using spark, hive, python, and really just about anything.☆20Feb 13, 2015Updated 11 years ago
- A type class for data of all sizes.☆15Jul 9, 2019Updated 7 years ago
- Collect local Mesos slave, underlying operating system and machine metrics and produce to Apache Kafka☆20Jan 29, 2016Updated 10 years ago
- Utilities for building distributed systems on top of mesos☆23Aug 25, 2018Updated 8 years ago
- A Real-Time Analytical Processing (RTAP) example using Spark/Shark☆51Feb 21, 2014Updated 12 years ago
- The code for the in memory data pipeline that was presented at Berlin Buzzwords 2015.☆10Jun 1, 2015Updated 11 years ago
- Practical utilities for spark applications☆11Feb 26, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Examples for Fast Data Processing with Spark☆59Sep 10, 2013Updated 12 years ago
- ☆16May 9, 2018Updated 8 years ago
- Hadoop MapReduce tool to convert Avro data files to Parquet format.☆32May 22, 2013Updated 13 years ago
- Sparse feature extraction with Spark☆30Jul 25, 2018Updated 8 years ago
- Low level integration of Spark and Kafka☆129Mar 15, 2018Updated 8 years ago
- An SBT plugin for automatically calling Avro code generation and a thin scala wrapper for reading and writing Avro files☆22Mar 8, 2018Updated 8 years ago
- ☆15Feb 26, 2013Updated 13 years ago
- Spark Mllib 1.6.0版本算法封装☆11Mar 8, 2017Updated 9 years ago
- Kiji BentoBox: Developer SDK for Kiji including a standalone zero-configuration HBase micro-cluster☆25Sep 26, 2014Updated 11 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Ambari YARN UTILS☆30Mar 30, 2023Updated 3 years ago
- A SQL-esque scripting language for spatial processing and ETL☆11Mar 4, 2019Updated 7 years ago
- This project contains the code to translate between Apache Spark and SFrame.☆20Jul 13, 2016Updated 10 years ago
- Open source formats for scalable genomic processing systems using Avro. Apache 2 licensed.☆43Feb 13, 2026Updated 6 months ago
- Simplifying robust end-to-end machine learning on Apache Spark.☆473Apr 18, 2017Updated 9 years ago
- Splash Project for parallel stochastic learning☆92Jun 16, 2017Updated 9 years ago
- Facebook social data modeling with Scala, HBase, and HPaste☆26Dec 7, 2015Updated 10 years ago