Data Sketches for Apache Spark
☆22Dec 22, 2022Updated 3 years ago
Alternatives and similar repositories for datasketches-spark
Users that are interested in datasketches-spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repo stores my Spark Tutorial slides.☆15Feb 8, 2016Updated 10 years ago
- Amundsen Gremlin☆22Updated this week
- A library for writing chemical and biological data management systems☆10Oct 24, 2019Updated 6 years ago
- A library to store metadata of relational databases including the schema, statistics, and integrity constraints.☆26Aug 7, 2018Updated 8 years ago
- High performance Privacy By Design using Matryoshka and Spark talk code☆13May 21, 2019Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ansible role for bootstrapping a server installation☆10Apr 12, 2022Updated 4 years ago
- Visualize column-level data lineage in Spark SQL☆92May 13, 2022Updated 4 years ago
- Deriving Spark DataFrame schemas from case classes☆44Jun 24, 2024Updated 2 years ago
- Parallel S3 bucket-bucket differential sync in Go☆42Dec 4, 2017Updated 8 years ago
- Get Twitter trends with twitter4j, stream it to a Kafka topic, save it to MongoDB and visualize in Google Maps☆13Sep 30, 2021Updated 4 years ago
- A simple kernel to interact with a redis database from IPython☆22Sep 25, 2017Updated 8 years ago
- Project for the talk on NLP using LSTM implementation from DL4J on Spark☆20May 6, 2016Updated 10 years ago
- A framework for PSL inference.☆22Nov 9, 2015Updated 10 years ago
- Concept Representation (Embedding) and Semantic Relatedness☆15Jul 3, 2019Updated 7 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆17Dec 20, 2022Updated 3 years ago
- A sql extension build on spark3 datasource v2 api, ex: hive v2 catalog support amoung multi clusters☆11May 7, 2022Updated 4 years ago
- An Erlang ingester for GreptimeDB, which is compatible with GreptimeDB protocol and lightweight.☆16Aug 20, 2026Updated 2 weeks ago
- ☆12Dec 7, 2021Updated 4 years ago
- Spark Monitoring☆14Feb 28, 2023Updated 3 years ago
- A quotation-based Scala DSL for scalable data analysis.☆65Jul 7, 2022Updated 4 years ago
- Blazing-fast query execution engine speaks Apache Spark language and has Arrow-DataFusion at its core.☆11Apr 23, 2022Updated 4 years ago
- DDSketch: A Fast and Fully-Mergeable Quantile Sketch with Relative-Error Guarantees.☆133Aug 31, 2026Updated last week
- PostgreSQL extension providing approximate algorithms based on apache/datasketches-cpp☆96Aug 9, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Golang JSON Specification API Inspired By JSH☆15Jul 19, 2016Updated 10 years ago
- 根据sql 文件生成javabean和数据模型啥的☆11Nov 23, 2025Updated 9 months ago
- Decentralized auction on top of ERGO.☆32May 17, 2023Updated 3 years ago
- Client libraries of end users of Apache Kyuubi☆11May 15, 2026Updated 3 months ago
- SCARFF (SCAlable Real-time Frauds Finder) is a framework which enables credit card fraud detection.☆20Feb 8, 2017Updated 9 years ago
- An online book.☆11Jan 24, 2015Updated 11 years ago
- scripts for testing TiDB☆11Feb 4, 2026Updated 7 months ago
- A library for Amazon Neptune that enables AWS Signature Version 4 signing for HTTP using Netty.☆18Aug 20, 2026Updated 2 weeks ago
- SHAPE/S∀F∃: static prover/type-checker for N-D array programming in Scala, a use case of intuitionistic type theory☆33Sep 9, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.☆16Jul 28, 2026Updated last month
- ☆16May 8, 2017Updated 9 years ago
- 百度云盘/百度网盘API接口上传文件,基于Golang1.18☆19Jan 10, 2023Updated 3 years ago
- Fundamentals of Apache Flink [video], published by Packt☆12Jan 30, 2023Updated 3 years ago
- Tasks API for Stateful Functions on Flink☆13Aug 22, 2026Updated 2 weeks ago
- ☆63Nov 6, 2022Updated 3 years ago
- The seamless translation layer☆25Feb 9, 2018Updated 8 years ago