Toolkit for Apache Spark ML for Feature clean-up, feature Importance calculation suite, Information Gain selection, Distributed SMOTE, Model selection and training, Hyper parameter optimization and selection, Model interprability.
☆191Jun 1, 2021Updated 5 years ago
Alternatives and similar repositories for automl-toolkit
Users that are interested in automl-toolkit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DEPRECATED: Integrating Jupyter with Databricks via SSH☆70Jun 28, 2022Updated 4 years ago
- Accelerator to rapidly deploy customized features for your business☆57Dec 10, 2023Updated 2 years ago
- Collection of Machine Learning Examples for Azure Databricks☆42Nov 11, 2020Updated 5 years ago
- Databricks Migration Tools☆43May 24, 2021Updated 5 years ago
- Grouped time series forecasting engine☆39Jun 23, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Capturing model drift and handling its response - Example webinar☆109Aug 1, 2019Updated 7 years ago
- Manage your Databricks deployments and CI with code.☆203Feb 28, 2023Updated 3 years ago
- ☆16Jun 12, 2020Updated 6 years ago
- SparklingGraph documentation☆10Jan 7, 2020Updated 6 years ago
- Data-driven software (python implementation)☆26Oct 20, 2023Updated 2 years ago
- [ARCHIVED] Moved to github.com/NVIDIA/spark-xgboost-examples☆72Jul 15, 2020Updated 6 years ago
- Koalas: pandas API on Apache Spark☆3,375Mar 20, 2024Updated 2 years ago
- DeltaOMS is a solution that help build a centralized repository of Delta Transaction logs and associated operational metrics/statistics f…☆42Nov 27, 2023Updated 2 years ago
- Joblib Apache Spark Backend☆248Mar 24, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Process, visualize and use data easily.☆20Jul 6, 2023Updated 3 years ago
- A simplified version of featuretools for Spark☆31Jun 14, 2019Updated 7 years ago
- ☆14Jun 13, 2024Updated 2 years ago
- Spark implementation of computing Shapley Values using monte-carlo approximation☆80Mar 20, 2023Updated 3 years ago
- TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows…☆2,276Jun 2, 2026Updated 3 months ago
- This repository contains the notebooks and presentations we use for our Databricks Tech Talks☆735Jan 6, 2025Updated last year
- Extensible Rules Engine for custom Dataframe / Dataset validation☆141May 7, 2024Updated 2 years ago
- Simplified custom plugins for Trino☆16Jul 29, 2024Updated 2 years ago
- MLBox is a powerful Automated Machine Learning python library.☆1,535Aug 6, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Distributed scikit-learn meta-estimators in PySpark☆287Apr 26, 2025Updated last year
- Distributed t-SNE via Apache Spark☆158Dec 9, 2017Updated 8 years ago
- Deep Learning Pipelines for Apache Spark☆1,988Mar 30, 2023Updated 3 years ago
- Geospatial clustering at massive scale☆112Mar 27, 2026Updated 5 months ago
- Hopsworks - Data-Intensive AI platform with a Feature Store☆1,307Feb 10, 2025Updated last year
- Databricks Terraform Provider☆600Updated this week
- Sample base images for Databricks Container Services☆223Jul 21, 2026Updated last month
- API for manipulating time series on top of Apache Spark: lagged time values, rolling statistics (mean, avg, sum, count, etc), AS OF joins…☆345Jul 10, 2026Updated 2 months ago
- Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark☆5,247Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Modular and minimalistic MLOps recipes☆71Mar 2, 2020Updated 6 years ago
- Example code for doing DataOps☆49Jan 26, 2021Updated 5 years ago
- Examples for Deep Learning/Feature Store/Spark/Flink/Hive/Kafka jobs and Jupyter notebooks on Hops☆116Jan 28, 2026Updated 7 months ago
- Version 1 of Technical Best Practices of Azure Databricks based on real world Customer and Technical SME inputs☆470Oct 6, 2023Updated 2 years ago
- An end-to-end Recommendation System built on Azure Databricks☆56Jul 29, 2019Updated 7 years ago
- My Study guide used to pass the CRT020 Spark Certification exam☆34Jan 6, 2020Updated 6 years ago
- A Python library for time series forecasting☆80Jun 11, 2025Updated last year