Official Dockerfile for Apache Spark
☆173Jul 23, 2026Updated 2 months ago
Alternatives and similar repositories for spark-docker
Users that are interested in spark-docker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Apache Spark Kubernetes Operator☆323Updated this week
- A Spark data source for reading Microsoft Excel files☆13Jul 1, 2024Updated 2 years ago
- A set of transformations for Kafka Connect☆23Mar 1, 2026Updated 6 months ago
- Kubernetes operator for managing the lifecycle of Apache Spark applications on Kubernetes.☆3,156Sep 18, 2026Updated last week
- Sparglim✨ makes PySpark App Configurable and Deploy Spark Connect Server Easier!☆42Jan 19, 2026Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A tool to get better debug info on spark's memory usage☆42Aug 21, 2019Updated 7 years ago
- Spark-Dashboard is an open-source monitoring solution for Apache Spark that provides real-time performance dashboards using containers an…☆140Sep 18, 2026Updated last week
- Python API for Deequ☆826Updated this week
- FITS data source for Spark SQL and DataFrames☆22Apr 12, 2023Updated 3 years ago
- Apache Spark Connect Client for Golang☆251May 15, 2026Updated 4 months ago
- Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.☆1,066Sep 15, 2026Updated last week
- Docker packaging for Apache Flink☆359Sep 18, 2026Updated last week
- ☆12Aug 11, 2026Updated last month
- Helm charts for Trino and Trino Gateway☆197Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A library for reading data from Amzon S3 with optimised listing using Amazon SQS using Spark SQL Streaming ( or Structured streaming).☆19Apr 20, 2024Updated 2 years ago
- ☆19May 7, 2026Updated 4 months ago
- Collection of NiFi-related stuff☆24Oct 27, 2022Updated 3 years ago
- Enables automatic refactoring and linting of Maven projects written in Scala using Scalafix.☆26Sep 15, 2026Updated last week
- Apache flink☆21Updated this week
- Helm Chart for deploying Spark history server in Amazon EKS for S3 Spark Event Logs☆31Apr 4, 2026Updated 5 months ago
- Testing Sandbox for Hadoop Ecosystem Components☆46Aug 20, 2026Updated last month
- Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.☆2,368Updated this week
- A simple spark standalone cluster for your testing environment purposses☆567Mar 6, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A JupyterHub authenticator using Kerberos☆12Sep 1, 2026Updated 3 weeks ago
- A re-implementation of Hadoop DistCP in Apache Spark☆47Dec 20, 2023Updated 2 years ago
- Kubernetes Helm Chart to deploy Apache Atlas☆16Oct 19, 2020Updated 5 years ago
- Apache DataFusion Comet Spark Accelerator☆1,286Updated this week
- A library that provides useful extensions to Apache Spark and PySpark.☆240Sep 10, 2026Updated 2 weeks ago
- Flux Operator Helm Charts☆26Updated this week
- How to setup a minimal Hadoop cluster using Docker☆10Mar 13, 2022Updated 4 years ago
- ☆12Jan 18, 2021Updated 5 years ago
- An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Tr…☆9,021Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Docker image for Spark history server on Kubernetes☆15Mar 13, 2020Updated 6 years ago
- Adapter for dbt that executes dbt pipelines on Apache Flink☆103Mar 19, 2024Updated 2 years ago
- simple inverted index full text search engine written in python☆13Oct 3, 2013Updated 12 years ago
- Gluten is a middle layer responsible for offloading JVM-based SQL engines' execution to native engines.☆1,602Updated this week
- Drop-in replacement for Apache Spark UI☆491Aug 20, 2026Updated last month
- Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.☆16Jul 28, 2026Updated 2 months ago
- Proxy for S3☆20Aug 20, 2026Updated last month