Launch AWS Elastic MapReduce jobs that process Common Crawl data.
☆49Feb 15, 2017Updated 9 years ago
Alternatives and similar repositories for elasticrawl
Users that are interested in elasticrawl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hadoop jobs for WikiReverse project. Parses Common Crawl data for links to Wikipedia articles.☆38Aug 12, 2018Updated 8 years ago
- [DEPRECATED] Landing page for cloud.gov. New repo: https://github.com/18F/cg-site☆12Nov 22, 2016Updated 9 years ago
- A simple Ruby example of how to process Common Crawl files using Elastic MapReduce☆29Mar 25, 2012Updated 14 years ago
- DKPro C4CorpusTools is a collection of tools for processing CommonCrawl corpus, including Creative Commons license detection, boilerplate…☆54Jun 12, 2020Updated 6 years ago
- ☆18Apr 29, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- fondz is a tool for auto-generating an "archival description" from a bag or series of bags.☆26Jun 30, 2014Updated 12 years ago
- Analyzing the April 2016 Data about the Usage of Sci-Hub☆28May 24, 2016Updated 10 years ago
- Fcrepo4 webapp plus optional fcrepo dependencies☆13Sep 30, 2020Updated 5 years ago
- Rails engine for working with storage of OpenAnnotations stored in Fedora4☆13Aug 4, 2016Updated 10 years ago
- Docker containers for running VIVO☆13Oct 26, 2016Updated 9 years ago
- Django app for managing PREMIS Events☆14Apr 28, 2026Updated 4 months ago
- Apache Nutch fork tunned for web services and data discovery.☆10May 18, 2015Updated 11 years ago
- Ansible deployment of fedora 4, single or clustered on ubuntu 14.04☆10Nov 11, 2015Updated 10 years ago
- Rewrite of original Salesforce Bulk API gem with full tests and API capability☆15Jul 14, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Common web archive utility code.☆65Sep 1, 2026Updated 3 weeks ago
- w3act is an annotation and curation tool for building web archive collections☆21Jan 30, 2024Updated 2 years ago
- Dynamic framework for Linked Data modelling, persistence, and full-text indexing.☆16Jan 13, 2020Updated 6 years ago
- A Ruby client for the OpenAI API support for multiple API configurations in a single app, robust and simple error handling, and network-l…☆20Feb 22, 2026Updated 7 months ago
- Set of scripts to aid in the download of the GDELT data files from www.gdeltproject.org☆12May 17, 2014Updated 12 years ago
- Curses-based audio/video visualization☆21Apr 4, 2013Updated 13 years ago
- A command-line tool for using CommonCrawl Index API at http://index.commoncrawl.org/☆203Oct 7, 2018Updated 7 years ago
- JavaScript library for getting geojson from the Wikipedia API☆22Sep 25, 2015Updated 10 years ago
- ☆14Sep 13, 2014Updated 12 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Templates for form letters to Canadian MPs☆20Jan 30, 2017Updated 9 years ago
- Collaborative collection development for web archives☆19Sep 5, 2019Updated 7 years ago
- GraphPass is a utility to filter networks and provide a default visualization output for Gephi or SigmaJS.☆17Nov 14, 2020Updated 5 years ago
- Events and Situations Ontology☆15Apr 20, 2018Updated 8 years ago
- Jekyll plugin to embed static IIIF images in jekyll pages☆23Mar 29, 2022Updated 4 years ago
- An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed…☆162Oct 8, 2025Updated 11 months ago
- A crawler based on Phantom. Allows discovery of dynamic content and supports custom scrapers.☆24May 31, 2017Updated 9 years ago
- Index URLs in Common Crawl☆197Sep 19, 2017Updated 9 years ago
- Ansible Roles and Playbooks for Princeton University Library☆18Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Fedora API Specification☆17May 6, 2021Updated 5 years ago
- A simple Node.js wrapper for the BitX API.☆11Jun 23, 2022Updated 4 years ago
- Various MARC command line utilities.☆37Apr 7, 2025Updated last year
- Notes to accompany Thomas Piketty, Capital in the Twenty-First Century (Harvard University Press, 2014)☆12May 10, 2020Updated 6 years ago
- ARK minter, binder, resolver☆24May 28, 2026Updated 3 months ago
- Python Linked Data Fragment Server.☆30Jun 4, 2018Updated 8 years ago
- A growing, online "textbook" for music theory and aural skills☆29Apr 27, 2017Updated 9 years ago