Apache-Spark based Data Flow(ETL) Framework which supports multiple read, write destinations of different types and also support multiple categories of transformation rules.
β26Jun 7, 2021Updated 5 years ago
Alternatives and similar repositories for DaFlow
Users that are interested in DaFlow are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simple Spark-powered ETL framework that just works πΊβ186Oct 2, 2025Updated 10 months ago
- An ETL framework in Scala for Data Engineersβ23Aug 30, 2022Updated 3 years ago
- Scala API for Apache Spark SQL high-order functionsβ15Aug 4, 2023Updated 3 years ago
- An extensible Scala framework for creating monitoring dashboards.β22Jan 12, 2023Updated 3 years ago
- Spark package to "plug" holes in data using SQL based rules β‘οΈ πβ28May 15, 2020Updated 6 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β24Apr 21, 2023Updated 3 years ago
- A dynamic data completeness and accuracy library at enterprise scale for Apache Sparkβ30May 13, 2026Updated 3 months ago
- Apache Parquet reader in Scala without Apache Spark - developed at Purdue Universityβ12Feb 17, 2017Updated 9 years ago
- Streaming Runtime Kubernetes' execution environment, designed to simplify the development and the operation of streaming data processing β¦β10Jan 18, 2023Updated 3 years ago
- β10Jun 5, 2021Updated 5 years ago
- A simplified, lightweight ETL Framework based on Apache Sparkβ588Aug 18, 2026Updated last week
- A framework for rapid reporting API development; with out of the box support for high cardinality dimension lookups with druid.β134Jan 17, 2025Updated last year
- Kafka Streams + Memcached (e.g. AWS ElasticCache) for low-latency in-memory lookupsβ13Nov 4, 2019Updated 6 years ago
- Source code for 'Real-Time Web Application Development' by Rami Vemulaβ12Dec 18, 2017Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Azure Cosmos DB - Custom Point in Time Restoreβ12Dec 7, 2022Updated 3 years ago
- A sink to save Spark Structured Streaming DataFrame into Hive tableβ23May 7, 2018Updated 8 years ago
- Parallel Streaming Transformation Loaderβ10Apr 23, 2019Updated 7 years ago
- HDInsight Developer Guideβ14Jun 27, 2018Updated 8 years ago
- β10Feb 18, 2021Updated 5 years ago
- Read-only primitive Java arrays backed by Direct Buffers and indexed using 64-bit indexesβ14Jun 8, 2017Updated 9 years ago
- Realtime social media data analytics with Apache Spark, Python, Kafka, Pandas, etcβ52Aug 25, 2016Updated 10 years ago
- Framework to make bots based on Microsoft Bot Framework.β13Oct 5, 2018Updated 7 years ago
- A solution for on-demand training and serving of Machine Learning models, using Azure Databricks and MLflowβ19Jul 17, 2020Updated 6 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Library for organizing batch processing pipelines in Apache Sparkβ43Jan 4, 2017Updated 9 years ago
- A highly available and infinitely scalable, drop-in replacement for Kafka Streamsβ21May 27, 2025Updated last year
- library for conducting propensity matching on spark scaleβ14Jun 27, 2023Updated 3 years ago
- This repository contains code for Spark Streamingβ26Mar 11, 2021Updated 5 years ago
- Play-Utils is a set of utilities for developing with Play Framework for Scala.β18Sep 25, 2018Updated 7 years ago
- Sample application describing a Bill of Materials scenario through an NPM (Node package manager) dependency explorer solution. Code by Chβ¦β16May 6, 2023Updated 3 years ago
- Streaming data changes to a Data Lake with Debezium and Delta LakeΒ pipelineβ77Feb 15, 2023Updated 3 years ago
- β13Jun 7, 2018Updated 8 years ago
- This extension for Visual Studio code enables you to click on Angular selectors in HTML files and be redirected to their definition in thβ¦β14Jul 21, 2018Updated 8 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Spark SQL index for Parquet tablesβ134May 6, 2021Updated 5 years ago
- Various data stream/batch process demo with Apache Scala Spark πβ12Feb 28, 2020Updated 6 years ago
- π¦ Dashboard to follow in real time the Covid-19 evolution.β14Jul 11, 2020Updated 6 years ago
- API gateway with Akka HTTPβ51Nov 27, 2020Updated 5 years ago
- Cloud workshop Azure ML 2020β12Dec 1, 2021Updated 4 years ago
- Maelstrom is an open source Kafka integration with Spark that is designed to be developer friendly, high performance (millisecond stream β¦β21Feb 6, 2017Updated 9 years ago
- Infuse AI into your application. Create and deploy a customer churn prediction model with IBM Cloud Private for Data, Db2 Warehouse, Sparβ¦β18Sep 17, 2025Updated 11 months ago