A simple Spark-powered ETL framework that just works πΊ
β186Oct 2, 2025Updated 10 months ago
Alternatives and similar repositories for setl
Users that are interested in setl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simplified, lightweight ETL Framework based on Apache Sparkβ588Jan 24, 2024Updated 2 years ago
- Generate fake data for Scala and Sparkβ15Dec 19, 2025Updated 7 months ago
- Apache-Spark based Data Flow(ETL) Framework which supports multiple read, write destinations of different types and also support multipleβ¦β26Jun 7, 2021Updated 5 years ago
- Smart Automation Tool for building modern Data Lakes and Data Pipelinesβ130Updated this week
- Qubole Sparklens tool for performance tuning Apache Sparkβ592Jun 26, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An ETL framework in Scala for Data Engineersβ23Aug 30, 2022Updated 3 years ago
- Extensible streaming ingestion pipeline on top of Apache Sparkβ47Jul 17, 2025Updated last year
- A home for LinkedIn's changes to Apache Icebergβ65Aug 3, 2026Updated last week
- A library enabling DAG structuring of data processing programs such as ETLsβ18Jul 19, 2026Updated 3 weeks ago
- Collection of open-source Spark tools & frameworks that have made the data engineering and data science teams at Swoop highly productiveβ191Oct 15, 2025Updated 9 months ago
- Essential Spark extensions and helper methods β¨π²β767Jun 22, 2026Updated last month
- A docker using the airflow with Hadoop ecosystem (hive, spark, and sqoop)β12May 2, 2021Updated 5 years ago
- Apache Spark based ETL Engineβ71Oct 18, 2016Updated 9 years ago
- Lab project to showcase Flink's performance differences between using a SQL query and implementing the same logic via the DataStream APIβ14Apr 15, 2020Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Basin is a visual programming editor for building Spark and PySpark pipelines. Easily build, debug, and deploy complex ETL pipelines fromβ¦β35Jan 5, 2023Updated 3 years ago
- Sample processing code using Spark 2.1+ and Scalaβ51Jun 28, 2020Updated 6 years ago
- Atomic Scala Book Solutions - for Beginners and first time Functional Programmersβ12Mar 10, 2020Updated 6 years ago
- An example pipeline that tests a Python project using pipenv for dependency management.β16Apr 14, 2026Updated 3 months ago
- A library that provides useful extensions to Apache Spark and PySpark.β240Updated this week
- Qubole Streaminglens tool for tuning Spark Structured Streaming Pipelinesβ17Jan 21, 2020Updated 6 years ago
- pyspark methods to enhance developer productivity π£ π― πβ687Jun 9, 2026Updated 2 months ago
- β64Nov 8, 2019Updated 6 years ago
- β24Apr 21, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Generate relevant synthetic data quickly for your projects. The Databricks Labs synthetic data generator (aka `dbldatagen`) may be used β¦β487Updated this week
- Multi-stage, config driven, SQL based ETL framework using PySparkβ26Sep 16, 2019Updated 6 years ago
- Waimak is an open-source framework that makes it easier to create complex data flows in Apache Spark.β76Apr 24, 2024Updated 2 years ago
- A dynamic data completeness and accuracy library at enterprise scale for Apache Sparkβ30May 13, 2026Updated 2 months ago
- A SparkSQL formatter based on https://github.com/zeroturnaround/sql-formatter, with customizations and extra features.β14Nov 7, 2024Updated last year
- π Docker image for AWS Glue Spark/Pythonβ23Sep 5, 2023Updated 2 years ago
- Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.β3,639Jul 21, 2026Updated 2 weeks ago
- A curated list of awesome Apache Spark packages and resources.β1,889Feb 27, 2026Updated 5 months ago
- A COBOL parser and Mainframe/EBCDIC data source for Apache Sparkβ170Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- API for manipulating time series on top of Apache Spark: lagged time values, rolling statistics (mean, avg, sum, count, etc), AS OF joinsβ¦β344Jul 10, 2026Updated last month
- Scalable CDC Pattern Implemented using PySparkβ18Oct 8, 2025Updated 10 months ago
- Data Lineage Tracking And Visualization Solutionβ665Updated this week
- Apache Spark testing helpers (dependency free & works with Scalatest, uTest, and MUnit)β457Apr 2, 2026Updated 4 months ago
- Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologiesβ¦β22Jan 10, 2019Updated 7 years ago
- The Internals of Spark SQLβ488Jan 25, 2026Updated 6 months ago
- Cloud based Data Platform based on Apache Sparkβ28Jun 30, 2026Updated last month