This package contains a generic implementation of greedy Information Theoretic Feature Selection (FS) methods. The implementation is based on the common theoretic framework presented by Gavin Brown. Implementations of mRMR, InfoGain, JMI and other commonly used FS filters are provided.
☆135May 5, 2022Updated 4 years ago
Alternatives and similar repositories for spark-infotheoretic-feature-selection
Users that are interested in spark-infotheoretic-feature-selection are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spark implementation of Fayyad's discretizer based on Minimum Description Length Principle (MDLP)☆43Jan 12, 2023Updated 3 years ago
- Featureselection methods as Spark MLlib Pipelines☆31Apr 29, 2018Updated 8 years ago
- Machine learning enhancements to Spark MlLib☆20Mar 19, 2015Updated 11 years ago
- Practice and Workshop on BigData and Cloud Computing using Docker Containers and OpenNebula. HDFS, hadoop and spark+R☆11Mar 16, 2017Updated 9 years ago
- Generic implementation of Information Theory-based Feature Selection methods. It also contains an Entropy Minimization Discretization imp…☆19Jul 21, 2014Updated 11 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Spark Extension : ML transformers, SQL aggregations, etc that are missing in Apache Spark☆145Jan 26, 2016Updated 10 years ago
- ☆25Mar 12, 2018Updated 8 years ago
- Distributed t-SNE via Apache Spark☆159Dec 9, 2017Updated 8 years ago
- Sparse feature extraction with Spark☆30Jul 25, 2018Updated 7 years ago
- Feature engineering toolkit for Spark MLlib.☆12Apr 1, 2017Updated 9 years ago
- Zen aims to provide the largest scale and the most efficient machine learning platform on top of Spark, including but not limited to logi…☆169Nov 17, 2018Updated 7 years ago
- This package contains the code for executing clustering validity indices in Spark. The package includes BD-Silhouette, BD-Dunn, Davies-Bo…☆10Oct 29, 2018Updated 7 years ago
- MLeap demo repository for use with MLeap blog posts☆11Jul 13, 2016Updated 9 years ago
- Old repo for R interface for GraphFrames☆13Mar 21, 2018Updated 8 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- dllib is a distributed deep learning library running on Apache Spark☆32Oct 26, 2017Updated 8 years ago
- Spark MLlib code optimized to efficiently support sparse data☆51Dec 22, 2016Updated 9 years ago
- PMML scoring library for Spark as SparkML Transformer☆21Oct 20, 2024Updated last year
- MLeap: Deploy ML Pipelines to Production☆1,538Mar 10, 2026Updated 3 months ago
- Guiones de prácticas para Álgebra Conmutativa y Computacional☆16Nov 9, 2023Updated 2 years ago
- Distributed DataFrame: Productivity = Power x Simplicity For Scientists & Engineers, on any Data Engine☆169Feb 26, 2021Updated 5 years ago
- A scalable machine learning library on Apache Spark☆797Aug 30, 2021Updated 4 years ago
- k-Nearest Neighbors algorithm on Spark☆241Nov 14, 2023Updated 2 years ago
- Recursos de Haskell☆18Mar 29, 2019Updated 7 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Experiments with scala native & libpcap☆10Mar 30, 2018Updated 8 years ago
- SparklingGraph provides easy to use set of features that will give you ability to proces large scala graphs using Spark and GraphX.☆154Jul 31, 2020Updated 5 years ago
- Affinity Propagation on Spark☆20May 31, 2021Updated 5 years ago
- Using JPMML Evaluator to validate the PMML models exported from Spark☆19May 1, 2017Updated 9 years ago
- Deeplearning framework running on Spark☆62Dec 16, 2023Updated 2 years ago
- Online Machine Learning Algorithms☆30Jun 14, 2023Updated 2 years ago
- Different entries to kaggle contests using Apache Spark☆13Jun 5, 2017Updated 9 years ago
- Hide udf☆18May 23, 2013Updated 13 years ago
- Model-based clustering package for mixed data☆13May 21, 2026Updated 3 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆11Dec 10, 2015Updated 10 years ago
- Trivial Spark app that counts Titan vertices☆10Mar 4, 2015Updated 11 years ago
- A web app for your finances☆12May 24, 2020Updated 6 years ago
- Sparkling Water provides H2O functionality inside Spark cluster☆978Nov 5, 2025Updated 7 months ago
- ☆78Oct 4, 2018Updated 7 years ago
- A Scala feature transformation library for data science and machine learning☆475Feb 7, 2025Updated last year
- Fitbit SDK example application.☆10Mar 7, 2023Updated 3 years ago