Spark pipelines that correspond to a series of Dataflow examples.
☆27May 5, 2019Updated 7 years ago
Alternatives and similar repositories for spark-examples
Users that are interested in spark-examples are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Google Cloud Dataflow pipelines such as Identity-By-State as well as useful utility classes.☆38Aug 9, 2023Updated 3 years ago
- Processing Logs at Scale using Cloud Dataflow☆60Mar 18, 2019Updated 7 years ago
- Data Science with Apache Spark and Spark Notebook☆30Jul 24, 2017Updated 9 years ago
- PostgreSQL and GreenPlum Data Source for Apache Spark☆35May 6, 2026Updated 3 months ago
- Java chat example app☆11Mar 11, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆16Jun 27, 2020Updated 6 years ago
- This repository is deprecated. All of its content and history has been moved to googleapis/google-cloud-node.☆13Jul 20, 2023Updated 3 years ago
- ScalaIO 2014 Workshop☆25Oct 23, 2014Updated 11 years ago
- A Gentle introduction to Machine Learning with Apache Spark☆11Mar 2, 2026Updated 5 months ago
- Google Cloud Dataflow provides a simple, powerful model for building both batch and streaming parallel data processing pipelines.☆163May 31, 2017Updated 9 years ago
- Optimizing downstream data processing with Amazon Kinesis Data Firehose and Amazon EMR running Apache Spark☆14Apr 14, 2023Updated 3 years ago
- Basic Spark utilities☆13Jul 24, 2026Updated 2 weeks ago
- DataSphere 产品文档☆12Sep 25, 2019Updated 6 years ago
- Fundamentals of Apache Flink [video], published by Packt☆12Jan 30, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Source of paper “A critique of the CAP theorem”☆16Dec 14, 2015Updated 10 years ago
- Docker-based utility for testing network failures and partitions in distributed applications☆10Oct 4, 2016Updated 9 years ago
- Sample how to use Camunda DMN decisions in a Zeebe Workflow☆10Apr 13, 2022Updated 4 years ago
- A set of widgets for Python's Orange Machine Learning to work with Apache Spark ML☆15Dec 24, 2016Updated 9 years ago
- Experimental: Multi-producer Single-consumer Queue☆12Jul 30, 2012Updated 14 years ago
- Akka Java cluster singleton example☆10Dec 5, 2023Updated 2 years ago
- ☆12Sep 22, 2023Updated 2 years ago
- Lab project to showcase Flink's performance differences between using a SQL query and implementing the same logic via the DataStream API☆14Apr 15, 2020Updated 6 years ago
- Tools for faster and optimized interaction with Teradata and large datasets.☆17Jul 11, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An interactive quiz application build on Kafka Streams, Spring MVC handler, and Vue☆12Aug 23, 2020Updated 5 years ago
- Repo for various Kubernetes applications☆18Dec 29, 2016Updated 9 years ago
- Google BigQuery support for Spark, Structured Streaming, SQL, and DataFrames with easy Databricks integration.☆70May 8, 2023Updated 3 years ago
- An example Akka project that runs in a cluster. This project is intended to be used to demonstrate Akka SBR.☆13Dec 5, 2023Updated 2 years ago
- Published data models for the Hercules vendor-agnostic SDN switch☆12Aug 15, 2020Updated 5 years ago
- This is a mirror of https://github.com/LucaCanali/sparkMeasure - sparkMeasure is a tool for performance troubleshooting of Apache Spark w…☆16May 21, 2026Updated 2 months ago
- A pyspark lib to validate data quality☆19Nov 11, 2022Updated 3 years ago
- API REST boilerplate using Spring Boot and Redis as database☆13Dec 26, 2018Updated 7 years ago
- ☆18Jul 5, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Due to lack of resources on how to deploy kafka with simple SASL authentication (just username and password) and how to write producer an…☆12Dec 29, 2021Updated 4 years ago
- Example to create lineage in Atlas with sqoop and spark☆14Apr 5, 2017Updated 9 years ago
- Exemplo de geração de dados (simulações) sobre contatos corporativos com a biblioteca Bogus em uma API REST criada com .NET 6 + ASP.NET C…☆11Jun 27, 2022Updated 4 years ago
- just some scripts that I use☆27Dec 19, 2012Updated 13 years ago
- Google Cloud Dataflow provides a simple, powerful model for building both batch and streaming parallel data processing pipelines.☆848Nov 25, 2020Updated 5 years ago
- A stateful serverless demo app running on AWS Lambda, using Apache Flink Stateful Functions☆15Oct 13, 2020Updated 5 years ago
- Word2Vec - Google's word2vec in Scala using UMASS factorie library for better hacking and research.☆16Apr 7, 2014Updated 12 years ago