Data for the HIPE 2022 shared task.
☆24May 15, 2026Updated 4 months ago
Alternatives and similar repositories for HIPE-2022-data
Users that are interested in HIPE-2022-data are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code and models for our CLEF-HIPE (Named Entity Processing on Historical Newspapers) submissions☆20Mar 27, 2023Updated 3 years ago
- Identifying Historical People, Places and other Entities: Shared Task on Named Entity Recognition and Linking on Historical Newspapers at…☆21Aug 1, 2024Updated 2 years ago
- Libraries, Archives and Museums (LAM)☆91Oct 4, 2022Updated 3 years ago
- CERberus -- guardian against character errors☆31Jul 3, 2026Updated 2 months ago
- Modules used for separating articles in (historical) newspapers and similar documents. This repository is part of the European Union's Ho…☆23Sep 2, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- LLM Benchmark Suite for Humanities Data☆26Updated this week
- Repository for "Towards Robust Named Entity Recognition for Historic German"☆18Dec 11, 2020Updated 5 years ago
- OCR post correction for old German corpus☆20Aug 29, 2022Updated 4 years ago
- ☆10Aug 5, 2019Updated 7 years ago
- Named Entity Recognition☆19Feb 13, 2026Updated 7 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ☆15Jul 11, 2022Updated 4 years ago
- Metrical position in Greek hexameter.☆13Sep 12, 2026Updated last week
- Pedalion trees☆13Jan 24, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Turn CTS TEI corpora into CEX collection files☆12Jun 16, 2021Updated 5 years ago
- A bunch of modules that use/extend CLTK in order to work with Greek and Latin corpora maintained by the Perseus DL☆13Oct 26, 2019Updated 6 years ago
- German GPT-2 model☆32Aug 17, 2021Updated 5 years ago
- Compiled tools, datasets, and other resources for historical text normalization.☆21Jun 18, 2019Updated 7 years ago
- ☆10May 8, 2026Updated 4 months ago
- BADLAD: Bengali Document Layout Analysis Dataset☆15Aug 11, 2026Updated last month
- ☆15Aug 22, 2026Updated 3 weeks ago
- Teaching materials for the Applied Data Analysis course at DHOxSS. Data science methods to analyse humanities data.☆41Jan 6, 2026Updated 8 months ago
- ☆20Feb 17, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- HuCit KB: a knowledge base of classical texts and citable text units.☆11Nov 17, 2021Updated 4 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆22Updated this week
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 9 months ago
- Codebase for running (conditional) probing experiments☆21Nov 13, 2022Updated 3 years ago
- Unofficial implementation of QaNER: Prompting Question Answering Models for Few-shot Named Entity Recognition.☆63Oct 15, 2022Updated 3 years ago
- Tutorial on NE processing for Digital Humanities - DH Utrech 2019☆24Jul 18, 2019Updated 7 years ago
- Patterns based on the W3C Web Annotation Model, primarily for use in linking resources describing historical phenomena with the places re…☆16Mar 6, 2020Updated 6 years ago
- [no more maintained] A free web widget to explore graphs by a simply and intuitive way!☆17May 4, 2010Updated 16 years ago
- omnesviae roman route planner☆16Jul 5, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- BERT and ELECTRA models trained on Europeana Newspapers☆39Dec 14, 2021Updated 4 years ago
- ☆14Aug 9, 2024Updated 2 years ago
- Archive of the XML files of the Mannheim / Heidelberg CAMENA Neo-Latin project☆20Oct 10, 2018Updated 7 years ago
- Public repository for Coptic SCRIPTORIUM Corpora Releases☆51Jul 27, 2026Updated last month
- ☆15Oct 21, 2023Updated 2 years ago
- Generating graph structures from OWL ontologies☆13Nov 21, 2017Updated 8 years ago
- Code for TACL 2020 paper "An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models"☆14Jul 31, 2020Updated 6 years ago