Example Repo to have full end to end pyspark testing via docker-compose
☆32Feb 6, 2023Updated 3 years ago
Alternatives and similar repositories for pyspark-testing-env
Users that are interested in pyspark-testing-env are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An example repo to demonstrate Django support in Pants☆41Feb 21, 2026Updated 7 months ago
- ☀️🦶 A lightweight framework for collaborative, open-source feature engineering☆33Oct 25, 2021Updated 4 years ago
- Data pipeline project using Data Factory, Databricks and Cosmosdb Graph, deployed using Azure DevOps, secured using firewalls and Azure A…☆11Dec 14, 2022Updated 3 years ago
- Implementation of Boundary Attributions for Normal (Vector) Explanations☆11Aug 13, 2021Updated 5 years ago
- In this article, you will learn how to set up a real-time data processing and analytics environment using Docker, MySQL, Redpanda, MinIO,…☆11Jun 27, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Repository for implementing alpha matting☆11Sep 4, 2019Updated 7 years ago
- Easily import a module and mock its dependencies in an isolated way.☆13May 19, 2022Updated 4 years ago
- Sets up the Databricks CLI in your GitHub Actions workflow.☆30Updated this week
- TensorFlow implementation of the "Prompt-to-Prompt Image Editing with Cross Attention Control" for Stable Diffusion☆15Mar 25, 2023Updated 3 years ago
- Deploy a scikit model using heroku and Flask☆15May 1, 2023Updated 3 years ago
- Match your fig size and font to conference formats.☆11Aug 16, 2021Updated 5 years ago
- Example project for building scalable data pipelines with Kedro and Ibis.☆14Dec 10, 2025Updated 9 months ago
- Turn browser clicks into reproducible scraping code.☆11Oct 27, 2024Updated last year
- A pyproject.toml conversion tool for Poetry to uv migration☆20Dec 28, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Instructions and code for the workshop "From Big Data to NLP Insights: Unlocking the Power of PySpark and Spark NLP"☆12May 9, 2023Updated 3 years ago
- Log all LSI MPT driver events to syslog☆16May 13, 2020Updated 6 years ago
- quadipy is a python package to help transform structured data into RDF graph format☆19Apr 14, 2023Updated 3 years ago
- HIVE: Evaluating the Human Interpretability of Visual Explanations (ECCV 2022)☆22Jan 19, 2023Updated 3 years ago
- A modern ELT demo using airbyte, dbt, snowflake and dagster☆29Nov 24, 2022Updated 3 years ago
- A Covid-19 data pipeline on AWS featuring PySpark/Glue, Docker, Great Expectations, Airflow, and Redshift, templated in CloudFormation an…☆24Nov 21, 2023Updated 2 years ago
- Repository of notebooks and related collateral used in the Databricks Demo Hub, showing how to use Databricks, Delta Lake, MLflow, and mo…☆26May 27, 2021Updated 5 years ago
- Enrolled in DataTalks Zoomcamp https://github.com/DataTalksClub/mlops-zoomcamp☆20Jun 27, 2022Updated 4 years ago
- A Terraform module to create and manage Identity and Access Management (IAM) Users on Amazon Web Services (AWS). https://aws.amazon.com/i…☆20Apr 6, 2022Updated 4 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆19May 22, 2024Updated 2 years ago
- The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for sever…☆292Aug 18, 2026Updated last month
- python and C++ projects - opencv☆12Nov 23, 2019Updated 6 years ago
- Nomad launcher/executor for Dagster☆22Aug 7, 2026Updated 2 months ago
- For a series of posts on Amazon MSK, Amazon EKS, and Amazon EMR☆68Jan 2, 2022Updated 4 years ago
- Fast Python dataclasses serialization☆14Oct 11, 2021Updated 4 years ago
- Samara changes data engineering by shifting from custom code to declarative configuration for complete ETL pipeline workflows.☆17Sep 12, 2026Updated 3 weeks ago
- A PoC script for adding dummy GitHub contributions to past dates☆12Nov 27, 2024Updated last year
- Making Databricks easy to use for R developers.☆26Oct 6, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Streamlit template for building SMART on FHIR apps in the Cerner ecosystem.☆11Sep 22, 2023Updated 3 years ago
- Making Time Speak! 🎙️☆29May 30, 2026Updated 4 months ago
- https://rogulski.it/blog/fastapi-async-db/☆11Jun 6, 2021Updated 5 years ago
- DuckDB Kernel - analytical execution runtime for Jupyter☆24Jun 29, 2026Updated 3 months ago
- For Udemy students: the official repository of Rock the JVM's Spark Streaming course☆26Jan 5, 2023Updated 3 years ago
- Demo converting streamlit uber nyc rides to use duckdb☆30Apr 9, 2023Updated 3 years ago
- Utility functions to support analytics over FHIR in BigQuery or Apache Spark☆15Jan 8, 2024Updated 2 years ago