CommonCrawl WARC/WET/WAT examples and processing code for Java + Hadoop
☆56Apr 26, 2021Updated 5 years ago
Alternatives and similar repositories for cc-warc-examples
Users that are interested in cc-warc-examples are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Sep 8, 2016Updated 9 years ago
- ☆23Feb 22, 2024Updated 2 years ago
- Tools to analyze web archives☆20Jul 12, 2016Updated 10 years ago
- The Historian's WARC Toolkit☆16May 14, 2015Updated 11 years ago
- A set of reusable Java components that implement functionality common to any web crawler☆259Jul 2, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- API to store Spring Batch (version 3+) job execution data in MongoDB☆12Jun 25, 2022Updated 4 years ago
- Connectivity with persistent store implementations☆12May 29, 2026Updated last month
- Scan a folder of document files of all types and extract the text into a CSV suitable for Overview☆26Mar 23, 2016Updated 10 years ago
- A repository with a useful structured 2007 SIC code table - As opposed to the less than useful HTML table provided by Companies House. If…☆17Mar 19, 2015Updated 11 years ago
- ☆10Mar 28, 2017Updated 9 years ago
- Getting started with Redis Streams & Java☆10Dec 2, 2024Updated last year
- Node.js library to copy files between several storage sources (ftp, http/https, s3, ssh, local, ...)☆11Oct 20, 2021Updated 4 years ago
- Docker Compose Prometheus with a Grafana UI☆10Apr 25, 2017Updated 9 years ago
- Index URLs in Common Crawl☆197Sep 19, 2017Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Explore GraphQL Subscriptions with a React Chat UI☆14Dec 15, 2023Updated 2 years ago
- A queue-controlled browser automation tool for improving web crawl quality☆68May 28, 2026Updated last month
- A library of examples showing how to use the Common Crawl corpus (2008-2012, ARC format)☆66Aug 5, 2016Updated 9 years ago
- Simple example on how to use Naive Bayes on Spark using the popular Reuters 21578 dataset☆23Jul 20, 2014Updated 12 years ago
- Crawlera tools☆26Feb 9, 2016Updated 10 years ago
- Library for Object Linking and Embedding (OLE) data types☆12Jun 24, 2026Updated 3 weeks ago
- ☆11Nov 11, 2016Updated 9 years ago
- ☆16May 10, 2023Updated 3 years ago
- Example Axon Framework based implementation based on CQRS Journey essay.☆16May 27, 2014Updated 12 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- This is the repository for the Redis Query Engine learning path on Redis University.☆22Apr 8, 2025Updated last year
- Converts WARC files to static HTML☆59Sep 18, 2025Updated 10 months ago
- Webscraping da Jurisprudência do Tribunal de Justiça do Distrito Federal☆11Dec 8, 2018Updated 7 years ago
- API implementation, User Interface, and more modules of the IPTC EXTRA project☆13Feb 14, 2022Updated 4 years ago
- Code for preservation simulation/modeling project☆10Aug 24, 2021Updated 4 years ago
- Vector Plugin for Solr: calculate dot product / cosine similarity on documents☆16Oct 24, 2018Updated 7 years ago
- Demonstrates how you can use Spring-MVC-Test-HTMLUnit with Cucumber-JVM☆18Aug 28, 2015Updated 10 years ago
- build a demo consul stack in vagrant with puppet and spring boot☆25Dec 15, 2014Updated 11 years ago
- Scripts and other tools for installing / operating Fusion on Kubernetes☆29Jul 8, 2026Updated 2 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- my dissertation!☆12Sep 6, 2022Updated 3 years ago
- Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.☆283Apr 25, 2018Updated 8 years ago
- Common Crawl One-Oh-One (aka "A Common Crawl Experiment")☆26Oct 31, 2014Updated 11 years ago
- ☆15Jan 3, 2015Updated 11 years ago
- Lint CLI for languagetool.☆14Apr 5, 2023Updated 3 years ago
- Metadata is lost when copying files around. It happens with cp, tar, rsync, Finder, Transmit, PathFinder. etc.☆27Jan 17, 2015Updated 11 years ago
- (deprecated) Please use new nlp4l instead.☆65Sep 22, 2016Updated 9 years ago