velib-v2: An ETL pipeline that employs batch and streaming jobs using Spark, Kafka, Airflow, and other tools, all orchestrated with Docker Compose.
☆20Aug 12, 2025Updated last year
Alternatives and similar repositories for end-to-end-etl-pipeline-jcdecaux-API
Users that are interested in end-to-end-etl-pipeline-jcdecaux-API are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Data Pipeline from the Global Historical Climatology Network DataSet☆27Dec 20, 2022Updated 3 years ago
- Public data and analytics for our open course☆35Mar 22, 2024Updated 2 years ago
- NoSQL extract, transform, load (ETL) toolkit with Python☆16Updated this week
- Sample project to demonstrate data engineering best practices☆227Feb 24, 2024Updated 2 years ago
- In this project I have built etl pipline which scraps the trending repository based on month,week and day LIVE extract other related info…☆12Sep 9, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A real-time reddit data streaming pipeline for sentiment analysis of various subreddits☆147Aug 23, 2023Updated 3 years ago
- Data pipeline that scrapes Rust cheater Steam profiles☆53Feb 13, 2022Updated 4 years ago
- Hands-on workshop with Apache Iceberg☆15Mar 13, 2024Updated 2 years ago
- ☆11Nov 18, 2022Updated 3 years ago
- StarCraft 2 Data Pipeline with Airflow, DuckDB and Streamlit☆17Mar 14, 2024Updated 2 years ago
- API/Data Platform for Ingesting, Storing, and Serving Data through Postgres, and Litestar☆11Sep 28, 2026Updated last week
- End to end data engineering project☆59Oct 27, 2022Updated 3 years ago
- Full-stack OpenTelemetry observability for Apache Spark☆20Updated this week
- An end-to-end workflow for processing streaming data on Azure.☆17Sep 20, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Docktor is a Web App that deploys an easy-to-use kit of analysis and scanning tools.☆13Nov 1, 2023Updated 2 years ago
- End to end data pipeline to extract and analyze submissions from any subreddit using Pushshift, python, dbt and BigQuery.☆12Jul 17, 2023Updated 3 years ago
- Steve's coffee shop recipe project for the Pluralsight Course "Git Fundamentals"☆21Mar 13, 2023Updated 3 years ago
- End-to-end ELT data engineering project☆23Dec 24, 2022Updated 3 years ago
- ☆16Jan 19, 2022Updated 4 years ago
- ecommerce GCP Streaming pipeline ― Cloud Storage, Compute Engine, Pub/Sub, Dataflow, Apache Beam, BigQuery and Tableau; GCP Batch pipelin…☆10Mar 9, 2022Updated 4 years ago
- Skooldio: Data Pipelines with Airflow☆23May 24, 2025Updated last year
- Instructions and code for the workshop "From Big Data to NLP Insights: Unlocking the Power of PySpark and Spark NLP"☆12May 9, 2023Updated 3 years ago
- Rock Solid Python with Type Hints Course Student Materials☆27Jul 8, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Sample Data Lakehouse deployed in Docker containers using Apache Iceberg, Minio, Trino and a Hive Metastore. Can be used for local testin…☆81Sep 2, 2023Updated 3 years ago
- Trino Iceberg Metadata Insights via Streamlit☆17Apr 9, 2025Updated last year
- rust-for-data☆53Sep 7, 2026Updated last month
- ☆14Jul 14, 2026Updated 2 months ago
- Pulsar Presto (outdated), go to https://github.com/apache/pulsar-sql instead☆18Oct 18, 2024Updated last year
- A benchmark for serverless analytic databases.☆26Jan 23, 2026Updated 8 months ago
- Model Context Protocol Server for the Observational Medical Outcomes Partnership (OMOP) Common Data Model☆29Jan 12, 2026Updated 8 months ago
- Spark-based pipeline to extract and parse monthly games from the Lichess database.☆22Sep 22, 2025Updated last year
- ☆19Jul 27, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DataTalks.Club's Data Engineering Zoomcamp Project☆25Jul 14, 2022Updated 4 years ago
- ZINDI GIZ NLP Agricultural Keyword Spotter 3rd place solution, Audio Classification☆11Sep 8, 2021Updated 5 years ago
- This is a demo streaming project simulating a music streaming service.☆33Aug 20, 2024Updated 2 years ago
- Docker with Airflow + Postgres + Spark cluster + JDK (spark-submit support) + Jupyter Notebooks☆24Apr 2, 2022Updated 4 years ago
- ☆25Dec 18, 2020Updated 5 years ago
- Open episode of the data engineering practice course☆31Jul 2, 2024Updated 2 years ago
- ☆13Feb 27, 2024Updated 2 years ago