Realtime Data Engineering Project
☆31Jan 12, 2025Updated last year
Alternatives and similar repositories for spark_clickhouse_streaming
Users that are interested in spark_clickhouse_streaming are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- code-snippets☆14Apr 9, 2026Updated 5 months ago
- This project provides an end-to-end data processing and visualization of visa numbers in Japan using PySpark and Plotly. The spark cluste…☆11Oct 11, 2023Updated 2 years ago
- ☆17Dec 30, 2020Updated 5 years ago
- ☆15Mar 29, 2024Updated 2 years ago
- ☆20Apr 3, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆10Aug 26, 2026Updated 3 weeks ago
- Applied Data Science training course (for updates and resources, read the ReadMe file below)☆16Sep 9, 2023Updated 3 years ago
- Retrieval Augmented Generation (RAG) on audio data with LangChain☆15Sep 26, 2023Updated 2 years ago
- ELT Data Pipeline implementation in Data Warehousing environment☆31May 2, 2025Updated last year
- The Kafka message scheduling tool.☆19Jan 20, 2025Updated last year
- Website for Applied Language Technology courses at the University of Helsinki☆19Aug 12, 2022Updated 4 years ago
- This is a demo project to compare two web scrapping frameworks, Playwright and Selenium and using the new Pipelining tool Dagster☆15Sep 9, 2021Updated 5 years ago
- Produce Kafka messages, consume them and upload into Cassandra, MongoDB.☆43Sep 26, 2023Updated 2 years ago
- Query Iceberg in Trino, Nessie as Catalog, and use minio to replace AWS S3☆27Aug 7, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This project leverages Hadoop, Spark, SQL, and Hive for efficient data integration, transformation, warehousing, and analytics. It provid…☆24Sep 30, 2023Updated 2 years ago
- ⚡ n8n on AWS EKS — Deploy n8n workflow automation on Amazon EKS☆16Jun 13, 2026Updated 3 months ago
- This project is about building a dimensional data warehouse in BigQuery by transforming an OLTP system to an OLAP system, using dbt as ou…☆16Dec 11, 2023Updated 2 years ago
- An end-to-end data pipeline for building Data Lake and supporting report using Apache Spark.☆16Jan 31, 2023Updated 3 years ago
- capstone project for Dataengineer.io bootcamp Public Repo☆12Feb 20, 2024Updated 2 years ago
- A Video Sharing Web Application Built Using Python3 and Django☆18Jan 2, 2022Updated 4 years ago
- ☆17Jul 28, 2026Updated last month
- ☆16Jun 7, 2024Updated 2 years ago
- ☆11Aug 20, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A fully integrated Django x Next.js tutorial on Google Docs☆16Mar 10, 2025Updated last year
- ☆21Nov 4, 2023Updated 2 years ago
- ☆13Mar 30, 2024Updated 2 years ago
- ☆16Aug 19, 2026Updated last month
- This project leverages Language Model (LLM) finetuning, Semantic Chunking, and a Retrieval-Augmented Generation (RAG) based Chatbot frame…☆19Jul 20, 2024Updated 2 years ago
- ☆22Sep 22, 2025Updated 11 months ago
- Ecommerce Realtime Data Pipeline (Data Modeling, Workflow Orchestration, Change Data Capture, Analytical Database and Dashboarding)☆72Mar 9, 2024Updated 2 years ago
- Mock streaming data generator☆18May 31, 2024Updated 2 years ago
- PySpark Tutorials and Materials☆19Mar 1, 2021Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Chroma maintenance CLI☆18Aug 15, 2025Updated last year
- Showcasing AI Data Modeling with Lineage & impact analysis on OpenMetadata☆20Apr 16, 2026Updated 5 months ago
- A Proxy service using FastAPI and Protocol Buffers (Proto3)☆13Jun 17, 2023Updated 3 years ago
- FLaNK AI Weekly covering Apache NiFi, Apache Flink, Apache Kafka, Apache Spark, Apache Iceberg, Apache Ozone, Apache Pulsar, and more...☆22Dec 29, 2025Updated 8 months ago
- Feature demos, integration guides & hands-on labs/projects using Kpow, Flex, Kafka, Flink, Iceberg & more☆53Jul 28, 2026Updated last month
- Glue ETL job or EMR Spark that gets from data catalog, modifies and uploads to S3 and Data Catalog☆13Aug 26, 2023Updated 3 years ago
- Local Environment to Practice Data Engineering☆143Dec 30, 2024Updated last year