A simple Spark-powered ETL framework that just works πΊ
β186Oct 2, 2025Updated 10 months ago
Alternatives and similar repositories for setl
Users that are interested in setl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simplified, lightweight ETL Framework based on Apache Sparkβ588Aug 18, 2026Updated last week
- Generate fake data for Scala and Sparkβ15Dec 19, 2025Updated 8 months ago
- Apache-Spark based Data Flow(ETL) Framework which supports multiple read, write destinations of different types and also support multipleβ¦β26Jun 7, 2021Updated 5 years ago
- Smart Automation Tool for building modern Data Lakes and Data Pipelinesβ130Updated this week
- Qubole Sparklens tool for performance tuning Apache Sparkβ593Jun 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An ETL framework in Scala for Data Engineersβ23Aug 30, 2022Updated 4 years ago
- Extensible streaming ingestion pipeline on top of Apache Sparkβ47Jul 17, 2025Updated last year
- A home for LinkedIn's changes to Apache Icebergβ66Updated this week
- A library enabling DAG structuring of data processing programs such as ETLsβ18Jul 19, 2026Updated last month
- Collection of open-source Spark tools & frameworks that have made the data engineering and data science teams at Swoop highly productiveβ191Oct 15, 2025Updated 10 months ago
- Essential Spark extensions and helper methods β¨π²β767Jun 22, 2026Updated 2 months ago
- Apache Spark based ETL Engineβ71Oct 18, 2016Updated 9 years ago
- Lab project to showcase Flink's performance differences between using a SQL query and implementing the same logic via the DataStream APIβ14Apr 15, 2020Updated 6 years ago
- Basin is a visual programming editor for building Spark and PySpark pipelines. Easily build, debug, and deploy complex ETL pipelines fromβ¦β35Jan 5, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Sample processing code using Spark 2.1+ and Scalaβ51Jun 28, 2020Updated 6 years ago
- Atomic Scala Book Solutions - for Beginners and first time Functional Programmersβ12Mar 10, 2020Updated 6 years ago
- A library that provides useful extensions to Apache Spark and PySpark.β240Aug 6, 2026Updated 3 weeks ago
- Qubole Streaminglens tool for tuning Spark Structured Streaming Pipelinesβ17Jan 21, 2020Updated 6 years ago
- pyspark methods to enhance developer productivity π£ π― πβ689Jun 9, 2026Updated 2 months ago
- β64Nov 8, 2019Updated 6 years ago
- A boilerplate project for Azure Big Data PaaS servicesβ14Dec 7, 2022Updated 3 years ago
- β24Apr 21, 2023Updated 3 years ago
- Generate relevant synthetic data quickly for your projects. The Databricks Labs synthetic data generator (aka `dbldatagen`) may be used β¦β491Aug 7, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Multi-stage, config driven, SQL based ETL framework using PySparkβ26Sep 16, 2019Updated 6 years ago
- Waimak is an open-source framework that makes it easier to create complex data flows in Apache Spark.β76Apr 24, 2024Updated 2 years ago
- The dbt-spark-livy adapter allows you to use dbt along with Apache Spark, by connecting via Apache Livyβ12Mar 30, 2023Updated 3 years ago
- A dynamic data completeness and accuracy library at enterprise scale for Apache Sparkβ30May 13, 2026Updated 3 months ago
- A SparkSQL formatter based on https://github.com/zeroturnaround/sql-formatter, with customizations and extra features.β14Nov 7, 2024Updated last year
- π Docker image for AWS Glue Spark/Pythonβ23Sep 5, 2023Updated 2 years ago
- Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.β3,642Jul 21, 2026Updated last month
- A curated list of awesome Apache Spark packages and resources.β1,893Feb 27, 2026Updated 6 months ago
- A COBOL parser and Mainframe/EBCDIC data source for Apache Sparkβ170Aug 8, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- API for manipulating time series on top of Apache Spark: lagged time values, rolling statistics (mean, avg, sum, count, etc), AS OF joinsβ¦β345Jul 10, 2026Updated last month
- Scalable CDC Pattern Implemented using PySparkβ18Oct 8, 2025Updated 10 months ago
- Data Lineage Tracking And Visualization Solutionβ667Aug 24, 2026Updated last week
- Apache Spark testing helpers (dependency free & works with Scalatest, uTest, and MUnit)β457Apr 2, 2026Updated 4 months ago
- Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologiesβ¦β22Jan 10, 2019Updated 7 years ago
- The Internals of Spark SQLβ490Jan 25, 2026Updated 7 months ago
- Cloud based Data Platform based on Apache Sparkβ28Jun 30, 2026Updated 2 months ago