This package contains a generic implementation of greedy Information Theoretic Feature Selection (FS) methods. The implementation is based on the common theoretic framework presented by Gavin Brown. Implementations of mRMR, InfoGain, JMI and other commonly used FS filters are provided.
☆134May 5, 2022Updated 4 years ago
Alternatives and similar repositories for spark-infotheoretic-feature-selection
Users that are interested in spark-infotheoretic-feature-selection are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spark implementation of Fayyad's discretizer based on Minimum Description Length Principle (MDLP)☆43Jan 12, 2023Updated 3 years ago
- Featureselection methods as Spark MLlib Pipelines☆30Apr 29, 2018Updated 8 years ago
- Machine learning enhancements to Spark MlLib☆20Mar 19, 2015Updated 11 years ago
- An improved implementation of the classical feature selection method: minimum Redundancy and Maximum Relevance (mRMR).☆82Apr 1, 2022Updated 4 years ago
- Generic implementation of Information Theory-based Feature Selection methods. It also contains an Entropy Minimization Discretization imp…☆19Jul 21, 2014Updated 12 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Spark Extension : ML transformers, SQL aggregations, etc that are missing in Apache Spark☆145Jan 26, 2016Updated 10 years ago
- ☆25Mar 12, 2018Updated 8 years ago
- Distributed t-SNE via Apache Spark☆158Dec 9, 2017Updated 8 years ago
- Feature engineering toolkit for Spark MLlib.☆12Apr 1, 2017Updated 9 years ago
- Sparse feature extraction with Spark☆30Jul 25, 2018Updated 8 years ago
- Spark-based GBM☆56Feb 19, 2020Updated 6 years ago
- This package contains the code for executing clustering validity indices in Spark. The package includes BD-Silhouette, BD-Dunn, Davies-Bo…☆10Oct 29, 2018Updated 7 years ago
- My MSc on Data Science final project. This is a library for Data Pre-processing Algorithms for Streaming in Flink (DPASF)☆18Jul 1, 2019Updated 7 years ago
- Zen aims to provide the largest scale and the most efficient machine learning platform on top of Spark, including but not limited to logi…☆169Nov 17, 2018Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- MLeap demo repository for use with MLeap blog posts☆11Jul 13, 2016Updated 10 years ago
- Old repo for R interface for GraphFrames☆13Mar 21, 2018Updated 8 years ago
- Model-based clustering package for mixed data☆13May 21, 2026Updated 2 months ago
- Example code for building your own MemSQL Streamliner Pipelines☆23Apr 18, 2017Updated 9 years ago
- PMML scoring library for Spark as SparkML Transformer☆21Oct 20, 2024Updated last year
- MLeap: Deploy ML Pipelines to Production☆1,540Jul 21, 2026Updated 3 weeks ago
- Distributed DataFrame: Productivity = Power x Simplicity For Scientists & Engineers, on any Data Engine☆169Feb 26, 2021Updated 5 years ago
- A scalable machine learning library on Apache Spark☆797Aug 30, 2021Updated 4 years ago
- Affinity Propagation on Spark☆20May 31, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- k-Nearest Neighbors algorithm on Spark☆241Nov 14, 2023Updated 2 years ago
- Experiments with scala native & libpcap☆10Mar 30, 2018Updated 8 years ago
- SparklingGraph provides easy to use set of features that will give you ability to proces large scala graphs using Spark and GraphX.☆154Jul 31, 2020Updated 6 years ago
- Using JPMML Evaluator to validate the PMML models exported from Spark☆19May 1, 2017Updated 9 years ago
- Online Machine Learning Algorithms☆30Jun 14, 2023Updated 3 years ago
- 主要解决ctr预估工程中的特征选择,特征编号(特征离散),单特征auc和logloss这3个问题.☆20Mar 30, 2017Updated 9 years ago
- Deeplearning framework running on Spark☆63Dec 16, 2023Updated 2 years ago
- Different entries to kaggle contests using Apache Spark☆13Jun 5, 2017Updated 9 years ago
- A implementation of the Self-Tuning Spectral Clustering algorithm, and more.☆12Sep 4, 2016Updated 9 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Deep Learning Pipelines for Apache Spark☆1,988Mar 30, 2023Updated 3 years ago
- Fast Correlation-Based Feature Selection☆32May 12, 2017Updated 9 years ago
- Feature Selection by Optimized LASSO algorithm☆17May 3, 2017Updated 9 years ago
- ☆78Oct 4, 2018Updated 7 years ago
- A library for time series analysis on Apache Spark☆1,197Oct 13, 2020Updated 5 years ago
- Apache Spark 2x Machine Learning Cookbook, published by Packt☆33Jul 23, 2025Updated last year
- [DEPRECATED] Tensorflow wrapper for DataFrames on Apache Spark☆744Jul 30, 2024Updated 2 years ago