Assembly of fundamental statistics implemented based on Apache Spark
☆31Feb 11, 2016Updated 10 years ago
Alternatives and similar repositories for StatisticsOnSpark
Users that are interested in StatisticsOnSpark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spark MLlib code optimized to efficiently support sparse data☆51Dec 22, 2016Updated 9 years ago
- Topic Modeling on Apache Spark☆94Mar 1, 2019Updated 7 years ago
- Yelp Restaurant Photo Classification - Kaggle competition☆11Apr 19, 2019Updated 7 years ago
- Parallel ML System - Bosen Java implementation☆27Jan 23, 2017Updated 9 years ago
- ☆21Dec 9, 2015Updated 10 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Step-by-step Deep Leaning Tutorials on Apache Spark using BigDL☆210Jan 3, 2023Updated 3 years ago
- Links to example code downloads for Learning Path: Get Started with Natural Language Processing Using Python, Spark, and Scala☆16Feb 23, 2017Updated 9 years ago
- Gradient Boosting Enhanced with Step-Wise Feature Augmentation☆17Jan 13, 2021Updated 5 years ago
- ☆20Dec 1, 2016Updated 9 years ago
- Machine learning evaluation database☆24Feb 7, 2018Updated 8 years ago
- Another, hopefully better, implementation of ALS on Spark☆14May 20, 2015Updated 11 years ago
- Tensorflow implementation of a Neural Attention Model for Abstractive Summarization.☆10Jul 20, 2020Updated 6 years ago
- Dropwizard Metrics reporter for Apache Spark☆28Dec 22, 2014Updated 11 years ago
- ☆62Jul 11, 2019Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- a spark custom window function example, to generate session IDs☆19Oct 26, 2017Updated 8 years ago
- Automatically exported from code.google.com/p/jbirch☆12Sep 6, 2022Updated 3 years ago
- Timeseries segmentation library☆12Mar 8, 2023Updated 3 years ago
- JVM related exercises☆11Jul 16, 2017Updated 9 years ago
- Parallel ML System - STRADS scheduler☆30Oct 4, 2018Updated 7 years ago
- Memory consumption estimator for Scala/Java☆27Nov 24, 2014Updated 11 years ago
- All presentations from Data Fest Kyiv 2017 http://datafest.in.ua☆13Apr 24, 2017Updated 9 years ago
- This project is for the notebooks, code, and data for the "Vocabulary Analysis of Job Descriptions" tutorial at PyData 2017 Seattle☆20Jul 12, 2017Updated 9 years ago
- A focused web crawler based on Playwright, RMQ, Kafka and Flink.☆14Feb 4, 2021Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- AugBoost: Gradient Boosting Enhanced with Step-Wise Feature Augmentation (2019 IJCAI paper)☆24Oct 22, 2019Updated 6 years ago
- A project with examples of using few commonly used data manipulation/processing/transformation APIs in Apache Spark 2.0.0☆26Aug 5, 2021Updated 5 years ago
- ☆21Feb 9, 2023Updated 3 years ago
- Manual for RStudio Server☆16Oct 5, 2013Updated 12 years ago
- Simple UDF to split JSON arrays into Hive arrays☆10Jun 24, 2016Updated 10 years ago
- Group project for the WorldQuant University module, risk management.☆13Feb 3, 2019Updated 7 years ago
- Python scripts to facilitate easy working☆11Mar 23, 2026Updated 5 months ago
- Word2Vec - Google's word2vec in Scala using UMASS factorie library for better hacking and research.☆16Apr 7, 2014Updated 12 years ago
- 『빅데이터 분석을 위한 스파크 2 프로그래밍』 예제 코드☆27Jun 25, 2017Updated 9 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Apache Flink CEP examples☆12Jul 22, 2016Updated 10 years ago
- 44th place solution in "Santander Customer Satisfaction"☆11May 16, 2016Updated 10 years ago
- Machine Learning with Scikit-Learn (material for pydata Amsterdam 2016)☆30Mar 21, 2016Updated 10 years ago
- keywords extraction☆17Dec 15, 2015Updated 10 years ago
- Workshop materials for scraping Twitter with Python☆13May 25, 2016Updated 10 years ago
- Building Annoy Index on Apache Spark☆72Jan 5, 2021Updated 5 years ago
- hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data☆14Jul 12, 2024Updated 2 years ago