A list of free datasets that provide streaming data
☆454May 29, 2026Updated 3 months ago
Alternatives and similar repositories for awesome-public-streaming-datasets
Users that are interested in awesome-public-streaming-datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Event data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic.☆101Jan 21, 2024Updated 2 years ago
- A list of publicly available datasets with real-time data maintained by the team at bytewax.io☆2,892Jul 10, 2026Updated last month
- Event data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic.☆547Jan 27, 2026Updated 7 months ago
- A data engineering project with Kafka, Spark Streaming, dbt, Docker, Airflow, Terraform, GCP and much more!☆902Apr 16, 2022Updated 4 years ago
- Stream processing pipeline from Finnhub websocket using Spark, Kafka, Kubernetes and more☆440Nov 28, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This repository contains the code for a realtime election voting system. The system is built using Python, Kafka, Spark Streaming, Postgr…☆49Dec 11, 2023Updated 2 years ago
- Este é um projeto de exemplo que demonstra um processo de ETL (Extração, Transformação e Carga) de dados usando Python, Polars e AWS Loca…☆15Sep 25, 2023Updated 2 years ago
- Sample code for building a Python application for Apache Flink on Kinesis Data Analytics.☆14Aug 30, 2023Updated 3 years ago
- A data pipeline with Kafka, Spark Streaming, dbt, Docker, Airflow, and GCP!☆12Jul 6, 2023Updated 3 years ago
- CLI tool to manage Kafka connectors☆10Mar 2, 2024Updated 2 years ago
- Hudi Demo Notebook☆11Mar 5, 2024Updated 2 years ago
- Kubernetes deployment of PrestoDB, Hive Metastore, and Minio S3-standard object store☆17Oct 20, 2022Updated 3 years ago
- ☆10Mar 12, 2021Updated 5 years ago
- Source code for the post, 'Getting Started with Data Analysis on AWS, using S3, Glue, Amazon Athena, and QuickSight'☆29Dec 22, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here 👇🏼☆45,083Updated this week
- End-to-End ELT data pipeline with Postgres, Airbyte, dbt, Dagster, Snowflake and Metabase☆12Jul 13, 2023Updated 3 years ago
- A topic-centric list of HQ open datasets.☆78,726Updated this week
- How to use Presto (with Hive metastore) and MinIO?☆29Mar 8, 2023Updated 3 years ago
- Project repository of Apache Airflow, deployed on Docker in Amazon EC2 via GitLab.☆16Sep 3, 2021Updated 4 years ago
- A kafka streams client library built on confluent-kafka-python☆66Sep 28, 2023Updated 2 years ago
- This repository contains an end-to-end data engineering project using Apache Flink, focused on performing sales analytics. The project de…☆12Nov 18, 2023Updated 2 years ago
- Code for my "Efficient Data Processing in SQL" book.☆63Aug 6, 2024Updated 2 years ago
- Generate relevant synthetic data quickly for your projects. The Databricks Labs synthetic data generator (aka `dbldatagen`) may be used …☆491Aug 7, 2026Updated 3 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A comprehensive Spark guide collated from multiple sources that can be referred to learn more about Spark or as an interview refresher.☆688Apr 22, 2022Updated 4 years ago
- A Redis clone written in Go☆43Aug 20, 2016Updated 10 years ago
- This project demonstrates how to use Apache Airflow to submit jobs to Apache spark cluster in different programming laguages using Python…☆48Mar 14, 2024Updated 2 years ago
- An Awesome List of Open-Source Data Engineering Projects☆3,272Oct 4, 2024Updated last year
- Data Pipeline from the Global Historical Climatology Network DataSet☆27Dec 20, 2022Updated 3 years ago
- A list of useful resources to learn Data Engineering from scratch☆4,010Jun 19, 2024Updated 2 years ago
- ☆23Jan 3, 2022Updated 4 years ago
- ☆40Updated this week
- akka http service for serving spark machine learning models☆15Aug 11, 2017Updated 9 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Data engineering with dbt, published by Packt☆110Sep 2, 2025Updated 11 months ago
- This construct builds some elements for you to quickly launch an EMR Serverless application. After submitting the Emr Serverless job, you…☆11Nov 18, 2025Updated 9 months ago
- Here lies all the pieces of portfolio projects and documents that I have been harvesting throughout the journey of learning Data Analysis…☆11Nov 22, 2023Updated 2 years ago
- Example end to end data engineering project.☆1,429Dec 8, 2022Updated 3 years ago
- ☆37Jun 3, 2023Updated 3 years ago
- Get data from API, run a scheduled script with Airflow, send data to Kafka and consume with Spark, then write to Cassandra☆145Jul 27, 2023Updated 3 years ago
- This is a basic Apache Pinot example for ingesting real-time MySQL change logs using Debezium☆28Jan 8, 2021Updated 5 years ago