Insight Data Engineering project: A platform built in HDFS, Spark and Airflow to help you to find social influencers from GitHub Network.
☆16May 21, 2024Updated 2 years ago
Alternatives and similar repositories for Git-Influencer
Users that are interested in Git-Influencer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SEJ Article notebooks☆16Nov 12, 2020Updated 5 years ago
- A collection of data engineering projects: data modeling, ETL pipelines, data lakes, infrastructure configuration on AWS, data warehousin…☆15Apr 29, 2021Updated 5 years ago
- Developed an ETL pipeline for a Data Lake that extracts data from S3, processes the data using Spark, and loads the data back into S3 as …☆17Oct 1, 2019Updated 6 years ago
- Tweepy Stream Example☆19Apr 23, 2019Updated 7 years ago
- Loan Default Prediction using PySpark, with jobs scheduled by Apache Airflow and Integration with Spark using Apache Livy☆22Dec 26, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Usage examples for byte-genie API☆12Apr 27, 2024Updated 2 years ago
- Use a AWS Glue Python Shell Job to connect to your Amazon Redshift cluster and execute a SQL script stored in Amazon S3.☆21Aug 8, 2022Updated 4 years ago
- This is a capstone project that entails building an end-to-end ETL (Extract-Transform-Load) Data pipeline which extracts UK accident and …☆18Jun 6, 2020Updated 6 years ago
- Infrastructure for researching self-driving databases☆33Jul 2, 2025Updated last year
- A production-grade data pipeline has been designed to automate the parsing of user search patterns to analyze user engagement. Extract d…☆24Nov 22, 2021Updated 4 years ago
- Built a stream processing data pipeline to get data from disparate systems into a dashboard using Kafka as an intermediary.☆29Aug 14, 2023Updated 3 years ago
- I am using confluent Kafka cluster to produce and consume scraped data. In this project, I've created a real-time data pipeline that uti…☆29May 2, 2023Updated 3 years ago
- JobAnalytics system consumes data from multiple sources and provides valuable information to both job hunters and recruiters.☆31Dec 8, 2022Updated 3 years ago
- Spark data pipeline that processes movie ratings data.☆31Aug 1, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A real-time streaming ETL pipeline for streaming and performing sentiment analysis on Twitter data using Apache Kafka, Apache Spark and D…☆29Aug 8, 2020Updated 6 years ago
- letter avatar is angular2 directive. It will generate avatar based on given text☆15Oct 31, 2019Updated 6 years ago
- Abandoned Object Dataset☆36Mar 24, 2015Updated 11 years ago
- This is python web scraper implemented using multithreading/multiprocessing/pool for amazon.com☆28Sep 23, 2019Updated 6 years ago
- This repository contains several example sub-projects related to data modeling using Redis with Redis OM for Python☆14Mar 2, 2022Updated 4 years ago
- ☆10Dec 22, 2018Updated 7 years ago
- Analyzing shifting trends in music through the ages☆13Mar 25, 2021Updated 5 years ago
- Uploads files with background uploads and progress feedback on modern browsers☆10Jun 3, 2026Updated 3 months ago
- general-purpose fast, stateless, and deterministic feature extractor written in golang for use in machine learning☆12Mar 17, 2018Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Mock async JavaScript libraries☆22Jan 2, 2016Updated 10 years ago
- Query Optimizer Service☆56Updated this week
- Resources, notebooks, assets for ML for Everyone Twitch stream☆14Jul 8, 2020Updated 6 years ago
- A collection of remark plugins used by HashiCorp to process markdown☆16Aug 22, 2025Updated last year
- Learn how to use Kinesis Firehose, AWS Glue, S3, and Amazon Athena by streaming and analyzing reddit comments in realtime. 100-200 level …☆45Apr 20, 2021Updated 5 years ago
- This repository contains code example in how to write search queries with OpenSearch Python client☆10Sep 20, 2023Updated 2 years ago
- ✋ Stop propagation for everyday events with Angular directives 🎩☆14Feb 4, 2018Updated 8 years ago
- Collection of Jupyter Notebooks in Python to monitor and improve your Watson Assistant workspaces☆10Jul 17, 2019Updated 7 years ago
- Tweets, Twitter, Search, Crawling, Tool, Big Data Analysis, Data Mining, Time Series Analysis☆15Jul 22, 2017Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ⚡️ A curated list of awesome things related to Infisical☆23Oct 16, 2023Updated 2 years ago
- Embed & Showcase your projects on websites☆17Jun 30, 2022Updated 4 years ago
- go-ima is a tool that checks if a file has been tampered with. It is useful in ensuring integrity in CI systems☆14Sep 28, 2023Updated 2 years ago
- PySpark functions and utilities with examples. Assists ETL process of data modeling☆103Dec 3, 2020Updated 5 years ago
- Qubole Streaminglens tool for tuning Spark Structured Streaming Pipelines☆17Jan 21, 2020Updated 6 years ago
- Semaphore demo CI/CD pipeline using Docker Compose and Python Flask☆13Jan 26, 2024Updated 2 years ago
- Simplified ETL process in Hadoop using Apache Spark. Has complete ETL pipeline for datalake. SparkSession extensions, DataFrame validatio…☆56May 6, 2023Updated 3 years ago