This is a capstone project that entails building an end-to-end ETL (Extract-Transform-Load) Data pipeline which extracts UK accident and traffic datasets from Amazon S3, clean and transform with Pyspark, transfer it back to S3 and finally load to Amazon Redshift (Distributed Database), from where the data can be queried for ad-hoc analyses.
☆18Jun 6, 2020Updated 6 years ago
Alternatives and similar repositories for UK_Accident_Traffic_ETL_Pipeline
Users that are interested in UK_Accident_Traffic_ETL_Pipeline are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- implementing an end-to-end tweets ETL/Analysis pipeline.☆59Dec 8, 2022Updated 3 years ago
- Spark + Python for Maketing Analytics☆10Apr 19, 2017Updated 9 years ago
- Example project for consuming AWS Kinesis streamming and save data on Amazon Redshift using Apache Spark☆11May 22, 2018Updated 8 years ago
- A summary of useful resources in order to learn about AI☆10May 9, 2020Updated 6 years ago
- Teaching notes from my Advanced SQL workshops as local lead instructor at General Assembly New York. The first edition was created for th…☆19Feb 14, 2020Updated 6 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- For this project I am creating an ETL (Extract, Transform, and Load) pipeline using Python, RegEx, and SQL Database. The goal is to retri…☆25Feb 9, 2021Updated 5 years ago
- Create a data pipeline on AWS to execute batch processing in a Spark cluster provisioned by Amazon EMR. ETL using managed airflow: extrac…☆10Jul 12, 2021Updated 5 years ago
- A repo to track data engineering projects☆14Nov 11, 2022Updated 3 years ago
- Insight Data Engineering project: A platform built in HDFS, Spark and Airflow to help you to find social influencers from GitHub Net…☆16May 21, 2024Updated 2 years ago
- Build Your Own Roadmap☆11Jul 8, 2020Updated 6 years ago
- 🏟☆28Nov 11, 2020Updated 5 years ago
- A collection of data engineering projects: data modeling, ETL pipelines, data lakes, infrastructure configuration on AWS, data warehousin…☆15Apr 29, 2021Updated 5 years ago
- Guide to CS Engineering and Interview Prep☆18Dec 26, 2024Updated last year
- An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.☆1,545Mar 9, 2020Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Tweepy Stream Example☆19Apr 23, 2019Updated 7 years ago
- Leetcode solution for weekly contest☆16Jan 11, 2020Updated 6 years ago
- Collection of notebooks☆17Oct 27, 2024Updated last year
- ☆16Mar 5, 2025Updated last year
- Apache Spark 2 for Beginners, published by Packt☆33Oct 31, 2022Updated 3 years ago
- Infrastructure for researching self-driving databases☆33Jul 2, 2025Updated last year
- data visualizations and R code for #TidyTuesday 2021☆16Feb 4, 2022Updated 4 years ago
- ☆14Aug 9, 2016Updated 10 years ago
- Exploring the use of options in creating small worlds for faster learning in RL Domains☆16Jan 23, 2012Updated 14 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- All Coding project for CS6515 GA☆17Jul 22, 2022Updated 4 years ago
- A complete example of an AWS Glue application that uses the Serverless Framework to deploy the infrastructure and DevContainers and/or Do…☆21Sep 11, 2026Updated 2 weeks ago
- A production-grade data pipeline has been designed to automate the parsing of user search patterns to analyze user engagement. Extract d…☆24Nov 22, 2021Updated 4 years ago
- Built a stream processing data pipeline to get data from disparate systems into a dashboard using Kafka as an intermediary.☆29Aug 14, 2023Updated 3 years ago
- An open-source repo to product management case studies.☆31Updated this week
- A real-time event pipeline around Kafka Ecosystem for Chicago Transit Authority.☆32Aug 14, 2023Updated 3 years ago
- Steven's 100DaysOfCloudRepo☆17Nov 22, 2020Updated 5 years ago
- Mastering Spark for Data Science, published by Packt☆51Apr 22, 2026Updated 5 months ago
- Spark data pipeline that processes movie ratings data.☆31Aug 1, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Apache Spark 3 for Data Engineering and Analytics with Python , By Packt publishing☆24Jul 23, 2023Updated 3 years ago
- A real-time streaming ETL pipeline for streaming and performing sentiment analysis on Twitter data using Apache Kafka, Apache Spark and D…☆29Aug 8, 2020Updated 6 years ago
- SageMaker specific extensions to TensorFlow.☆54Jul 23, 2024Updated 2 years ago
- These are the Jupyter notebooks for the Big Data specialization in the Data Science Program.☆15Apr 3, 2020Updated 6 years ago
- My Data Engineer Capstone project. A consolidated dataset with several jobs around the world.☆13May 22, 2023Updated 3 years ago
- Project files for the post: Running PySpark Applications on Amazon EMR using Apache Airflow: Using the new Amazon Managed Workflows for A…☆41Jul 6, 2022Updated 4 years ago
- in this prj we will find and detect the lane and we are able to find the area that the car should be in☆24Mar 28, 2022Updated 4 years ago