Koalas: pandas API on Apache Spark
☆3,372Mar 20, 2024Updated 2 years ago
Alternatives and similar repositories for koalas
Users that are interested in koalas are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Tr…☆9,040Updated this week
- The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, a…☆28,314Updated this week
- Modin: Scale your Pandas workflows by changing a single line of code☆10,393Feb 10, 2026Updated 7 months ago
- Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per s…☆8,513Apr 1, 2026Updated 6 months ago
- Deep Learning Pipelines for Apache Spark☆1,987Mar 30, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Parallel computing with task scheduling☆13,932Sep 29, 2026Updated last week
- Always know what to expect from your data.☆11,866Updated this week
- Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.☆3,648Sep 16, 2026Updated 3 weeks ago
- Jupyter magics and kernels for working with remote Spark clusters☆1,367Sep 9, 2025Updated last year
- An open source python library for automated feature engineering☆7,687Sep 11, 2026Updated 3 weeks ago
- Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark☆5,246Updated this week
- Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet f…☆1,895Jan 2, 2026Updated 9 months ago
- 📚 Parameterize, execute, and analyze notebooks☆6,490Updated this week
- Build, Manage and Deploy AI/ML Systems☆10,297Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Apache Spark - A unified analytics engine for large-scale data processing☆44,147Updated this week
- Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interacting…☆4,782Sep 10, 2026Updated 3 weeks ago
- Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.☆43,992Updated this week
- The Open Source Feature Store for AI/ML☆7,323Updated this week
- Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics☆17,190Updated this week
- Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark☆1,535Dec 2, 2024Updated last year
- MLeap: Deploy ML Pipelines to Production☆1,546Updated this week
- 1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.☆13,717Sep 11, 2026Updated 3 weeks ago
- Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and…☆11,015Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner☆2,632Mar 20, 2024Updated 2 years ago
- A game theoretic approach to explain the output of any machine learning model.☆25,798Updated this week
- the portable Python dataframe library☆6,676Updated this week
- A better notebook for Scala (and more)☆4,596Jan 27, 2026Updated 8 months ago
- 🦉 Data Versioning and ML Experiments☆15,909Updated this week
- Apache Airflow - A platform to programmatically author, schedule, and monitor workflows☆47,119Updated this week
- cuDF - GPU DataFrame Library☆9,772Updated this week
- State of the Art Natural Language Processing☆4,158Sep 30, 2026Updated last week
- Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.☆14,680Sep 17, 2026Updated 3 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- (Legacy) Command Line Interface for Databricks☆396Oct 5, 2023Updated 3 years ago
- Voilà turns Jupyter notebooks into standalone web applications☆5,949Updated this week
- A Python package for manipulating 2-dimensional tabular data structures☆1,877Sep 1, 2026Updated last month
- Fit interpretable models. Explain blackbox machine learning.☆6,956Updated this week
- TensorFlowOnSpark brings TensorFlow programs to Apache Spark clusters.☆3,846Jul 10, 2023Updated 3 years ago
- Prefect is a workflow orchestration framework for building resilient data pipelines in Python.☆23,989Updated this week
- A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rew…☆2,169May 19, 2026Updated 4 months ago