Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published 24 days ago • 3
Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections Paper • 2608.18957 • Published 24 days ago • 1
Institutional Books Collection A growing corpus of public domain books from library collections, seeded by Harvard Library. • 12 items • Updated 22 days ago • 9
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published 24 days ago • 12
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published 24 days ago • 12
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 10 days ago • 7
institutional/institutional-newspapers-crop-classifier-text-model2vec Text Classification • 32.4M • Updated 9 days ago • 26 • 1
Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections Paper • 2608.18957 • Published 24 days ago • 1
Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published 24 days ago • 3
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published 24 days ago • 12
Institutional Books Collection A growing corpus of public domain books from library collections, seeded by Harvard Library. • 12 items • Updated 22 days ago • 9
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 10 days ago • 7