How to build an awesome data engineering team
☆101Sep 11, 2019Updated 7 years ago
Alternatives and similar repositories for data-engineering
Users that are interested in data-engineering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Learning from multiple companies in Silicon Valley. Netflix, Facebook, Google, Startups☆896May 8, 2022Updated 4 years ago
- A list of useful resources to learn Data Engineering from scratch☆4,032Updated this week
- ☆13Oct 6, 2019Updated 7 years ago
- Few projects related to Data Engineering including Data Modeling, Infrastructure setup on cloud, Data Warehousing and Data Lake developme…☆1,997Aug 26, 2022Updated 4 years ago
- Sharing interesting and noteworthy Data Engineering content☆68Oct 21, 2016Updated 9 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A curated list of data engineering tools for software developers☆9,160Sep 7, 2026Updated last month
- This project contains the code to translate between Apache Spark and SFrame.☆20Jul 13, 2016Updated 10 years ago
- Selected resources for SRE/DevOps professionals covering various Computer Science areas: Software Engineering & Architecture, Operations,…☆29Jan 4, 2018Updated 8 years ago
- ☆12Apr 16, 2016Updated 10 years ago
- A boilerplate project for Azure Big Data PaaS services☆14Dec 7, 2022Updated 3 years ago
- Simple demonstration of how to build a complex real time machine learning visualization tool.☆16Mar 26, 2016Updated 10 years ago
- Coursera, Big Data Essentials: HDFS, MapReduce and Spark RDD☆12Jun 18, 2019Updated 7 years ago
- 神棚(kamidana) is command line jinja2 template☆11Sep 23, 2026Updated 2 weeks ago
- ☆196Feb 25, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is an example api which covers some topics about api creation with dotnet-core2☆10Dec 8, 2022Updated 3 years ago
- Big Data Engineering practice project, including ETL with Airflow and Spark using AWS S3 and EMR☆91Jul 17, 2019Updated 7 years ago
- An experimental tool to synchronize source Databricks deployment with a target Databricks deployment.☆50Jan 21, 2024Updated 2 years ago
- RAG applications repo for Uplimit course☆10Jul 20, 2025Updated last year
- ☆16Jun 25, 2019Updated 7 years ago
- Projects done in the Data Engineer Nanodegree Program by Udacity.com☆179Dec 8, 2022Updated 3 years ago
- Repo to migrate old wiki to, esp for devs and code examples☆183Oct 18, 2016Updated 9 years ago
- ☆13Jun 14, 2017Updated 9 years ago
- Browser Automation with Python and Selenium by Packt Publishing☆11Jan 30, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Amazon Redshift offers a common query interface against data stored in fast, local storage as well as data from high-capacity, inexpensiv…☆13Nov 26, 2018Updated 7 years ago
- PySpark for Beginners by Packt Pyblishing☆15Jan 30, 2023Updated 3 years ago
- ☆24Aug 12, 2023Updated 3 years ago
- ☆45Apr 21, 2022Updated 4 years ago
- A Python rookie looking to learn Python by analysing data from PC game Football Manager 2020☆11Sep 12, 2020Updated 6 years ago
- An Awesome List of Open-Source Data Engineering Projects☆3,301Oct 4, 2024Updated 2 years ago
- Python for Agricultural Engineers☆13Feb 5, 2016Updated 10 years ago
- Simple system monitor for Ubuntu and Nginx. It uses Node.js and internal commands to retrieve the information.☆10Sep 24, 2015Updated 11 years ago
- In this workshop you will launch an Amazon Redshift cluster in your AWS account and load sample data ~ 100GB using TPCH dataset. You will…☆24Nov 28, 2018Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Feb 20, 2020Updated 6 years ago
- FakeNilc is a set of tools written in python3 to train Machine Learning models for fake news detection. Models trained using this mini-fr…☆18Jan 21, 2022Updated 4 years ago
- Pydantic model for the AsyncAPI (v2) specification schema☆14Apr 9, 2026Updated 6 months ago
- Kubeflow for Poets: A Guide to Containerization of the Machine Learning Production Pipeline☆14Mar 28, 2019Updated 7 years ago
- Example end to end data engineering project.☆1,439Dec 8, 2022Updated 3 years ago
- Reference Architectures for Datalakes on AWS☆77May 13, 2020Updated 6 years ago
- Polls sample app for IBM BlueMix PaaS☆14Jun 25, 2014Updated 12 years ago