Agile Data Preparation Workflows madeΒ easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
β1,536Dec 2, 2024Updated last year
Alternatives and similar repositories for optimus
Users that are interested in optimus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π A spreadsheet-like data preparation web app that works over Optimus (Pandas, Dask, cuDF, Dask-cuDF, Spark and Vaex)β141Jul 15, 2023Updated 3 years ago
- Koalas: pandas API on Apache Sparkβ3,372Mar 20, 2024Updated 2 years ago
- Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per sβ¦β8,510Apr 1, 2026Updated 3 months ago
- NumPy and Pandas interface to Big Dataβ3,189Sep 29, 2023Updated 2 years ago
- Modin: Scale your Pandas workflows by changing a single line of codeβ10,393Feb 10, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- An open source python library for automated feature engineeringβ7,666Updated this week
- pyspark methods to enhance developer productivity π£ π― πβ687Jun 9, 2026Updated last month
- Always know what to expect from your data.β11,675Updated this week
- 1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.β13,654Apr 22, 2026Updated 3 months ago
- A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewβ¦β2,170May 19, 2026Updated 2 months ago
- Parallel computing with task schedulingβ13,871Jul 20, 2026Updated last week
- Open-source low code data preparation library in python. Collect, clean and visualization your data in python with a few lines of code.β2,247Jun 27, 2024Updated 2 years ago
- Build, Manage and Deploy AI/ML Systemsβ10,198Updated this week
- Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering andβ¦β10,937Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- π Parameterize, execute, and analyze notebooksβ6,462Jul 6, 2026Updated 3 weeks ago
- An orchestration platform for the development, production, and observation of data assets.β15,909Updated this week
- A Python package for manipulating 2-dimensional tabular data structuresβ1,876Updated this week
- the portable Python dataframe libraryβ6,612Updated this week
- A light-weight, flexible, and expressive statistical data testing libraryβ4,411Jul 18, 2026Updated last week
- Clean APIs for data cleaning. Python implementation of R package Janitorβ1,501Jul 20, 2026Updated last week
- Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.β3,638Updated this week
- The Open Source Feature Store for AI/MLβ7,178Updated this week
- Prefect is a workflow orchestration framework for building resilient data pipelines in Python.β23,494Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- π¦ Data Versioning and ML Experimentsβ15,776Jul 21, 2026Updated last week
- Python Helper library for Jupyter Notebooksβ1,040Feb 16, 2021Updated 5 years ago
- A Python Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming.β10,050Sep 11, 2025Updated 10 months ago
- Jupyter magics and kernels for working with remote Spark clustersβ1,364Sep 9, 2025Updated 10 months ago
- The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, aβ¦β27,235Updated this week
- Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visβ¦β18,752Jul 18, 2026Updated last week
- Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interactingβ¦β4,780Jul 1, 2026Updated 3 weeks ago
- Easy pipelines for pandas DataFrames.β729Jul 6, 2026Updated 3 weeks ago
- A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflowβ2,089Dec 15, 2023Updated 2 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- MLeap: Deploy ML Pipelines to Productionβ1,539Updated this week
- A package which efficiently applies any function to a pandas dataframe or series in the fastest available mannerβ2,639Mar 20, 2024Updated 2 years ago
- Visual analysis and diagnostic tools to facilitate machine learning model selection.β4,400Feb 19, 2025Updated last year
- A library for debugging/inspecting machine learning classifiers and explaining their predictionsβ2,777Apr 8, 2026Updated 3 months ago
- Engine for AI/ML/Data tracking, visualization, explainability, drift detection, and dashboards for Polyaxon.β534Jun 17, 2026Updated last month
- re_data - fix data issues before your users & CEO would discover them πβ1,566Apr 30, 2024Updated 2 years ago
- Automatic extraction of relevant features from time series:β9,277Jul 6, 2026Updated 3 weeks ago