This repository will help you to learn about databricks concept with the help of examples. It will include all the important topics which we need in our real life experience as a data engineer. We will be using pyspark & sparksql for the development. At the end of the course we also cover few case studies.
☆105Sep 26, 2025Updated 10 months ago
Alternatives and similar repositories for ApacheSpark
Users that are interested in ApacheSpark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Airflow & DBT Cloud Integrated Project Presented at Lagos DBT Community Meetup & DataFestAfrica 23☆13Oct 11, 2023Updated 2 years ago
- ETL (Extract, Transform and Load) with the Spark Python API (PySpark) and Hadoop Distributed File System (HDFS)☆17Dec 18, 2018Updated 7 years ago
- Leetcode SQL Solutions☆196Aug 26, 2023Updated 2 years ago
- Simple ETL pipeline using Python☆29May 22, 2023Updated 3 years ago
- PySpark Cheatsheet☆36Jan 18, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Jan 17, 2022Updated 4 years ago
- Jupyter Notebook showing how to process Telecom datasets using PySpark (SparkSQL and DataFrames) and plotting the results using Matplotli…☆17Dec 3, 2018Updated 7 years ago
- Ravi Azure ADB ADF Repository☆65Jan 25, 2025Updated last year
- Collection of Databricks and Jupyter Notebooks☆22Feb 9, 2026Updated 6 months ago
- Simplified ETL process in Hadoop using Apache Spark. Has complete ETL pipeline for datalake. SparkSession extensions, DataFrame validatio…☆56May 6, 2023Updated 3 years ago
- PySpark Cheat Sheet - example code to help you learn PySpark and develop apps faster☆498Oct 15, 2024Updated last year
- This repository contains code for Spark Streaming☆26Mar 11, 2021Updated 5 years ago
- A shell script to automate the operations of sqoop☆11Mar 29, 2021Updated 5 years ago
- Basic framework utilities to quickly start writing production ready Apache Spark applications☆36Dec 15, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Solution to all projects of Udacity's Data Engineering Nanodegree: Data Modeling with Postgres & Cassandra, Data Warehouse with Redshift,…☆58Oct 20, 2022Updated 3 years ago
- Basin is a visual programming editor for building Spark and PySpark pipelines. Easily build, debug, and deploy complex ETL pipelines from…☆35Jan 5, 2023Updated 3 years ago
- Develop ML models predict taxi trip duration in NYC. Ranked : Top 6% | RMSLE : 0.377 (Kaggle) | #DS☆17Jan 7, 2023Updated 3 years ago
- ☆28Jun 14, 2022Updated 4 years ago
- This data project can be used as a take-home assignment to learn Pyspark and Data Engineering.☆20Feb 19, 2023Updated 3 years ago
- This repo contains commands that data engineers use in day to day work.☆64Feb 4, 2023Updated 3 years ago
- Resources and projects from Udacity Data Engineering with AWS nano degree programme☆30Apr 12, 2023Updated 3 years ago
- Case Study's from Danny Ma's Serious SQL Course☆19Aug 4, 2022Updated 4 years ago
- Template for Scala Spark with Unit Test☆13Jul 24, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Apache Spark Interview Question and Answers☆21Oct 13, 2020Updated 5 years ago
- Databricks Platform - Architecture, Security, Automation and much more!!☆57Aug 6, 2026Updated last week
- This repo is mostly created for pyspark and hive related interview questions.☆64Jan 6, 2026Updated 7 months ago
- Dockerizing an Apache Spark Standalone Cluster☆42Jun 29, 2022Updated 4 years ago
- An end-to-end data engineering pipeline to create a dashboard for the latest content on the r/Stocks subreddit☆20Aug 5, 2022Updated 4 years ago
- End to end data engineering project☆59Oct 27, 2022Updated 3 years ago
- Complete SQL Project for data analysis with source code.☆376Oct 11, 2022Updated 3 years ago
- A four-day course on Python, the Scientific Python stack and PySpark, adapted from a training course I gave to one of our clients in Dece…☆10Feb 3, 2016Updated 10 years ago
- In this project, we will build and ETL(Extract,Transform,Load) pipeline using the Spotify API on AWS. The pipeline will retrieve data fro…☆25May 6, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- adidas Data Mesh implementation☆12May 13, 2022Updated 4 years ago
- (Python, PySpark)☆10Nov 15, 2020Updated 5 years ago
- ☆95Sep 14, 2022Updated 3 years ago
- Series follows learning from Apache Spark (PySpark) with quick tips and workaround for daily problems in hand☆55Sep 30, 2023Updated 2 years ago
- Mastering AWS CloudFormation Second Edition, published by packt☆18Oct 23, 2023Updated 2 years ago
- With everything I learned from DEZoomcamp from datatalks.club, this project performs a batch processing on AWS for the cycling dataset wh…☆15Jan 4, 2026Updated 7 months ago
- A project with examples of using few commonly used data manipulation/processing/transformation APIs in Apache Spark 2.0.0☆26Aug 5, 2021Updated 5 years ago