This is a capstone project that entails building an end-to-end ETL (Extract-Transform-Load) Data pipeline which extracts UK accident and traffic datasets from Amazon S3, clean and transform with Pyspark, transfer it back to S3 and finally load to Amazon Redshift (Distributed Database), from where the data can be queried for ad-hoc analyses.
☆18Jun 6, 2020Updated 6 years ago
Alternatives and similar repositories for UK_Accident_Traffic_ETL_Pipeline
Users that are interested in UK_Accident_Traffic_ETL_Pipeline are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This projects entails performing in-depth descriptive analysis and data visualization on United Kingdom Road Traffic and Accident dataset…☆11Jul 20, 2020Updated 6 years ago
- implementing an end-to-end tweets ETL/Analysis pipeline.☆59Dec 8, 2022Updated 3 years ago
- Spark + Python for Maketing Analytics☆10Apr 19, 2017Updated 9 years ago
- Primer curso de Craftech Academy - Marzo 2021☆12Aug 3, 2021Updated 5 years ago
- Example project for consuming AWS Kinesis streamming and save data on Amazon Redshift using Apache Spark☆11May 22, 2018Updated 8 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A summary of useful resources in order to learn about AI☆10May 9, 2020Updated 6 years ago
- Teaching notes from my Advanced SQL workshops as local lead instructor at General Assembly New York. The first edition was created for th…☆19Feb 14, 2020Updated 6 years ago
- Road Traffic Accident prediction using Machine Learning Classification Technique☆10Nov 16, 2020Updated 5 years ago
- Create a data pipeline on AWS to execute batch processing in a Spark cluster provisioned by Amazon EMR. ETL using managed airflow: extrac…☆10Jul 12, 2021Updated 5 years ago
- ☆17Sep 23, 2021Updated 4 years ago
- Data Analysis of Road Traffic Accidents to Minimize the rate of accidents☆15May 18, 2018Updated 8 years ago
- Udemy実践Pythonデータサイエンスの資料ページ☆16Nov 9, 2024Updated last year
- A repo to track data engineering projects☆14Nov 11, 2022Updated 3 years ago
- This is a multiclass classification project to classify severity of road accidents into three categories. this project is based on real-w…☆21Jul 10, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Insight Data Engineering project: A platform built in HDFS, Spark and Airflow to help you to find social influencers from GitHub Net…☆16May 21, 2024Updated 2 years ago
- Data mining algorithms with Python☆10Jun 26, 2019Updated 7 years ago
- A collection of data engineering projects: data modeling, ETL pipelines, data lakes, infrastructure configuration on AWS, data warehousin…☆15Apr 29, 2021Updated 5 years ago
- A Data Visualization project on the French traffic accidents database☆19Aug 27, 2019Updated 7 years ago
- Self-study plan to achieve mastery in data science☆30Mar 29, 2023Updated 3 years ago
- Roadmap to becoming a web developer in 2017 in spanish, Roadmap para ser un desarrollador web en el 2017☆15Jun 16, 2017Updated 9 years ago
- Repository for Traffic Accident Benchmark for Causality Recognition (ECCV 2020)☆34Jun 30, 2021Updated 5 years ago
- IV 2020 "CSG: Critical Scenario Generation from Real Traffic Accidents"☆22Jun 8, 2026Updated 2 months ago
- Developed an ETL pipeline for a Data Lake that extracts data from S3, processes the data using Spark, and loads the data back into S3 as …☆17Oct 1, 2019Updated 6 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- a convenient way to anonymize your data for analytics☆22Nov 7, 2021Updated 4 years ago
- An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.☆1,542Mar 9, 2020Updated 6 years ago
- Tweepy Stream Example☆19Apr 23, 2019Updated 7 years ago
- 使用aiohttp+asyncio简易的上海链家租房爬虫☆25Jan 21, 2017Updated 9 years ago
- Loan Default Prediction using PySpark, with jobs scheduled by Apache Airflow and Integration with Spark using Apache Livy☆22Dec 26, 2020Updated 5 years ago
- Samples of ML models learning from source code☆20Nov 28, 2022Updated 3 years ago
- Collection of notebooks☆17Oct 27, 2024Updated last year
- ☆16Mar 5, 2025Updated last year
- Apache Spark 2 for Beginners, published by Packt☆33Oct 31, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Infrastructure for researching self-driving databases☆33Jul 2, 2025Updated last year
- data visualizations and R code for #TidyTuesday 2021☆16Feb 4, 2022Updated 4 years ago
- ☆14Aug 9, 2016Updated 10 years ago
- A complete example of an AWS Glue application that uses the Serverless Framework to deploy the infrastructure and DevContainers and/or Do…☆21Updated this week
- A production-grade data pipeline has been designed to automate the parsing of user search patterns to analyze user engagement. Extract d…☆24Nov 22, 2021Updated 4 years ago
- Built a stream processing data pipeline to get data from disparate systems into a dashboard using Kafka as an intermediary.☆29Aug 14, 2023Updated 3 years ago
- Demonstration of using Apache Spark to build robust ETL pipelines while taking advantage of open source, general purpose cluster computin…☆24Aug 11, 2023Updated 3 years ago