Some class materials for a data processing course using PySpark
☆53Dec 3, 2022Updated 3 years ago
Alternatives and similar repositories for data_processing_course
Users that are interested in data_processing_course are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Create LAMP Stack using terraform with AWS☆11Feb 15, 2023Updated 3 years ago
- Hadoop Examples☆10Jul 1, 2022Updated 4 years ago
- Add gevent support to DataStax Python Driver for Apache Cassandra☆11Jun 10, 2020Updated 6 years ago
- ☆11Dec 14, 2015Updated 10 years ago
- Projects from my Hadoop training sessions☆16Feb 22, 2018Updated 8 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Ansible Playbook to create LAMP in CentOS 7 with Apache, MySQL, PHP.☆10Dec 28, 2018Updated 7 years ago
- All Certification and preparation, examples & others☆11Oct 18, 2018Updated 7 years ago
- ☆14Aug 24, 2021Updated 4 years ago
- Automated (Ansible) installation of HDP via Ambari Blueprint☆16Mar 10, 2017Updated 9 years ago
- Analytics projects using Big Data eco-systems (Hadoop, Spark, Storm)☆17Dec 27, 2021Updated 4 years ago
- Projects from Udacity Data Streaming Nanodegree☆15Aug 14, 2023Updated 2 years ago
- Ansible playbooks for Apache Spark on kube☆27Jul 20, 2017Updated 9 years ago
- Testbench for experimenting with Apache Hive at any data scale.☆64Jul 10, 2017Updated 9 years ago
- Python API for Informatica PowerCenter (pmrep, pmcmd)☆21Sep 17, 2017Updated 8 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- PySpark Tutorial for Beginners on Google Colab: Hands-On Guide☆17Sep 13, 2020Updated 5 years ago
- Python and C++ implementation of the problems from Clean Code Handbook - LeetCode 50 Common Interview Questions☆33Feb 6, 2023Updated 3 years ago
- How to create the Invoice Template Design In HTML and CSS☆11Jan 20, 2022Updated 4 years ago
- Set of Shell scripts to automate Linux from Scratch, based on the book 7.8☆31Jan 10, 2018Updated 8 years ago
- Finance 🏦 Data Builder 🛠️ @ postgres 🐘☆22Feb 11, 2021Updated 5 years ago
- Udacity Data Engineering Nanodegree Projects☆11Sep 5, 2019Updated 6 years ago
- Apache Spark (Scala, PySpark, SparkR) Code, Tricks, and References☆69Jan 21, 2019Updated 7 years ago
- DevOps☆16May 17, 2021Updated 5 years ago
- A content-based recommender system for books using the Project Gutenberg text corpus☆29Feb 20, 2017Updated 9 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Examples To Help You Learn Apache Spark☆76Oct 8, 2018Updated 7 years ago
- My applied big data analytic project with pyspark.☆10Sep 21, 2022Updated 3 years ago
- Lab environment based on vagrant to learn ex200/ex300 rhcsa/rhce☆39Mar 16, 2017Updated 9 years ago
- Using PyRaider You can scan installed dependencies known security vulnerabilities. It uses publicly known exploits, vulnerabilities datab…☆18May 18, 2022Updated 4 years ago
- A collection of data analysis projects done using PySpark via Jupyter notebooks.☆10Oct 8, 2022Updated 3 years ago
- Benchmarks to compare golang geohash implementations☆12Aug 6, 2018Updated 7 years ago
- Optimal cache stampede prevention☆16May 11, 2017Updated 9 years ago
- Spark 2.0 Python Machine Learning examples☆99Oct 7, 2019Updated 6 years ago
- Sentiment Analysis of a Twitter Topic with Spark Structured Streaming☆54Dec 12, 2018Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ansible roles to install an Spark Standalone cluster (HDFS/Spark/Jupyter Notebook) or Ambari based Spark cluster☆62Jan 30, 2024Updated 2 years ago
- A boilerplate for writing PySpark Jobs☆393Jan 21, 2024Updated 2 years ago
- AlvinToh Learning Repository for The Ultimate Hands-On Hadoop - Tame your Big Data!☆10May 23, 2018Updated 8 years ago
- My Machine Learning & Deep Learning Papers Notes.☆11Jul 17, 2018Updated 8 years ago
- ecommerce GCP Streaming pipeline ― Cloud Storage, Compute Engine, Pub/Sub, Dataflow, Apache Beam, BigQuery and Tableau; GCP Batch pipelin…☆11Mar 9, 2022Updated 4 years ago
- Methods for the parallel and distributed analysis and mining of the Protein Data Bank using MMTF and Apache Spark.☆68Mar 27, 2023Updated 3 years ago
- running apache spark with docker swarm☆34Feb 25, 2021Updated 5 years ago