A simple Spark-powered ETL framework that just works πΊ
β186Oct 2, 2025Updated 11 months ago
Alternatives and similar repositories for setl
Users that are interested in setl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simplified, lightweight ETL Framework based on Apache Sparkβ589Sep 1, 2026Updated 2 weeks ago
- Generate fake data for Scala and Sparkβ15Dec 19, 2025Updated 9 months ago
- Apache-Spark based Data Flow(ETL) Framework which supports multiple read, write destinations of different types and also support multipleβ¦β26Jun 7, 2021Updated 5 years ago
- Smart Automation Tool for building modern Data Lakes and Data Pipelinesβ130Updated this week
- Qubole Sparklens tool for performance tuning Apache Sparkβ593Jun 26, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- An ETL framework in Scala for Data Engineersβ23Aug 30, 2022Updated 4 years ago
- Extensible streaming ingestion pipeline on top of Apache Sparkβ47Jul 17, 2025Updated last year
- A home for LinkedIn's changes to Apache Icebergβ67Aug 28, 2026Updated 3 weeks ago
- A library enabling DAG structuring of data processing programs such as ETLsβ18Jul 19, 2026Updated 2 months ago
- Collection of open-source Spark tools & frameworks that have made the data engineering and data science teams at Swoop highly productiveβ192Oct 15, 2025Updated 11 months ago
- Essential Spark extensions and helper methods β¨π²β767Jun 22, 2026Updated 2 months ago
- A docker using the airflow with Hadoop ecosystem (hive, spark, and sqoop)β12May 2, 2021Updated 5 years ago
- Apache Spark based ETL Engineβ71Oct 18, 2016Updated 9 years ago
- Lab project to showcase Flink's performance differences between using a SQL query and implementing the same logic via the DataStream APIβ14Apr 15, 2020Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Basin is a visual programming editor for building Spark and PySpark pipelines. Easily build, debug, and deploy complex ETL pipelines fromβ¦β35Jan 5, 2023Updated 3 years ago
- Sample processing code using Spark 2.1+ and Scalaβ51Jun 28, 2020Updated 6 years ago
- Atomic Scala Book Solutions - for Beginners and first time Functional Programmersβ12Mar 10, 2020Updated 6 years ago
- An example pipeline that tests a Python project using pipenv for dependency management.β16Apr 14, 2026Updated 5 months ago
- A library that provides useful extensions to Apache Spark and PySpark.β240Sep 10, 2026Updated last week
- Qubole Streaminglens tool for tuning Spark Structured Streaming Pipelinesβ17Jan 21, 2020Updated 6 years ago
- pyspark methods to enhance developer productivity π£ π― πβ689Jun 9, 2026Updated 3 months ago
- β64Nov 8, 2019Updated 6 years ago
- A boilerplate project for Azure Big Data PaaS servicesβ14Dec 7, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β24Apr 21, 2023Updated 3 years ago
- Multi-stage, config driven, SQL based ETL framework using PySparkβ26Sep 16, 2019Updated 7 years ago
- Waimak is an open-source framework that makes it easier to create complex data flows in Apache Spark.β76Apr 24, 2024Updated 2 years ago
- The dbt-spark-livy adapter allows you to use dbt along with Apache Spark, by connecting via Apache Livyβ12Mar 30, 2023Updated 3 years ago
- A dynamic data completeness and accuracy library at enterprise scale for Apache Sparkβ30May 13, 2026Updated 4 months ago
- A SparkSQL formatter based on https://github.com/zeroturnaround/sql-formatter, with customizations and extra features.β14Nov 7, 2024Updated last year
- Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.β3,646Updated this week
- A curated list of awesome Apache Spark packages and resources.β1,898Feb 27, 2026Updated 6 months ago
- A COBOL parser and Mainframe/EBCDIC data source for Apache Sparkβ170Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Scalable CDC Pattern Implemented using PySparkβ18Oct 8, 2025Updated 11 months ago
- API for manipulating time series on top of Apache Spark: lagged time values, rolling statistics (mean, avg, sum, count, etc), AS OF joinsβ¦β345Jul 10, 2026Updated 2 months ago
- Data Lineage Tracking And Visualization Solutionβ667Sep 10, 2026Updated last week
- Apache Spark testing helpers (dependency free & works with Scalatest, uTest, and MUnit)β457Apr 2, 2026Updated 5 months ago
- Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologiesβ¦β22Jan 10, 2019Updated 7 years ago
- The Internals of Spark SQLβ491Jan 25, 2026Updated 7 months ago
- Cloud based Data Platform based on Apache Sparkβ28Jun 30, 2026Updated 2 months ago