Spark data pipeline that processes movie ratings data.
☆31Aug 1, 2026Updated last month
Alternatives and similar repositories for spark-movies-etl
Users that are interested in spark-movies-etl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Create a data pipeline on AWS to execute batch processing in a Spark cluster provisioned by Amazon EMR. ETL using managed airflow: extrac…☆10Jul 12, 2021Updated 5 years ago
- ☆25Dec 18, 2020Updated 5 years ago
- Various data stream/batch process demo with Apache Scala Spark 🚀☆12Feb 28, 2020Updated 6 years ago
- A real-time streaming ETL pipeline for streaming and performing sentiment analysis on Twitter data using Apache Kafka, Apache Spark and D…☆29Aug 8, 2020Updated 6 years ago
- PySpark functions and utilities with examples. Assists ETL process of data modeling☆103Dec 3, 2020Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- NoSQL extract, transform, load (ETL) toolkit with Python☆16Updated this week
- Example project for consuming AWS Kinesis streamming and save data on Amazon Redshift using Apache Spark☆11May 22, 2018Updated 8 years ago
- In this project I have built etl pipline which scraps the trending repository based on month,week and day LIVE extract other related info…☆12Sep 9, 2023Updated 3 years ago
- 🌟 An end-to-end full-stack data science project, including modelling, MLOps, and data storytelling. ✨☆16Aug 30, 2025Updated last year
- Boilerplate for PySpark on Cloud Kubernetes☆33Oct 12, 2021Updated 4 years ago
- Surface crack images classification using PyTorch Lightning☆11Jun 17, 2020Updated 6 years ago
- Power Pop Health is a collection of content intended to simplify the process of ingesting and prepping Healthcare Open Data using Azure d…☆18May 23, 2022Updated 4 years ago
- Insight Data Engineering project: A platform built in HDFS, Spark and Airflow to help you to find social influencers from GitHub Net…☆16May 21, 2024Updated 2 years ago
- StarCraft 2 Data Pipeline with Airflow, DuckDB and Streamlit☆17Mar 14, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code for youtube channel☆10Apr 15, 2022Updated 4 years ago
- Apache NiFi deployment on OpenShift☆13Jul 18, 2023Updated 3 years ago
- A Data Engineering Project that implements an ETL data pipeline using Dagster, Apache Spark, Streamlit, MinIO, Metabase, Dbt, Polars, Doc…☆25Nov 19, 2024Updated last year
- A collection of data engineering projects: data modeling, ETL pipelines, data lakes, infrastructure configuration on AWS, data warehousin…☆15Apr 29, 2021Updated 5 years ago
- This is a simple ETL using Airflow. First, we fetch data from API (extract). Then, we drop unused columns, convert to CSV, and validate (…☆24Oct 12, 2019Updated 6 years ago
- Command line client for the Fugue API☆14Mar 7, 2023Updated 3 years ago
- Starter application demonstrating how to connect a NestJS API to a PlanetScale MySQL database☆11May 6, 2026Updated 4 months ago
- dbt project for the domestic heating agent-based model at Centre for Net Zero.☆12Nov 15, 2022Updated 3 years ago
- Powershell Scripts for Power BI☆13Sep 20, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Apr 13, 2022Updated 4 years ago
- The supply chain dataset presents a comprehensive set of information related to product sales, manufacturing, and logistics. The challeng…☆24Sep 30, 2023Updated 2 years ago
- A custom AWS credential provider that allows your Hadoop or Spark application access S3 file system by assuming a role☆10Jan 9, 2026Updated 8 months ago
- Basic Udacity project using pandas for bikeshare data exploration☆10Nov 29, 2021Updated 4 years ago
- Tweepy Stream Example☆19Apr 23, 2019Updated 7 years ago
- Loan Default Prediction using PySpark, with jobs scheduled by Apache Airflow and Integration with Spark using Apache Livy☆22Dec 26, 2020Updated 5 years ago
- Generate OpenAPI 3.x.x using Pydantic☆11Feb 9, 2023Updated 3 years ago
- ☆10Feb 12, 2026Updated 7 months ago
- Set of scripts to terminate various GCP resources to save cash and cats 🐈☆14Jan 25, 2021Updated 5 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The goal of this project is to track the expenses of Uber Rides and Uber Eats through data Engineering processes using technologies such …☆124Jun 29, 2022Updated 4 years ago
- Code for my Medium article: "How you can quickly deploy your ML models with FastAPI"☆12Mar 18, 2021Updated 5 years ago
- This is a capstone project that entails building an end-to-end ETL (Extract-Transform-Load) Data pipeline which extracts UK accident and …☆18Jun 6, 2020Updated 6 years ago
- Educational project on how to build an ETL (Extract, Transform, Load) data pipeline, orchestrated with Airflow.☆354Jan 12, 2022Updated 4 years ago
- Cost effective data pipelines code repository☆16Sep 9, 2023Updated 3 years ago
- Delta Sharing + MLflow for ML model & experiment exchange (arcuate delta - a fan shaped river delta)☆22Jan 29, 2026Updated 7 months ago
- Postgres Node.js Express TypeScript application boilerplate with best practices for API development.☆13Jul 9, 2022Updated 4 years ago