

dat-019
Historical Newspaper OCR Dataset
Text strings mapped to bounding boxes on images — Download scans from library archives; extract text with Tesseract; manually correct errors. Stored as JSON / XML / PNG.
OCRPythonText DataResearch
What you get
- Source Code
- Documentation
- PPT
- Dataset Sample
Technical Details
Data Type
Text strings mapped to bounding boxes on images
Hardware / Tools
Python + Tesseract OCR + Public-domain newspaper scans
Collection Method
Download scans from library archives; extract text with Tesseract; manually correct errors
Storage Format
JSON / XML / PNG
How it works.
Common questions about ordering, delivery, and support.
More from Data Collection.
View all
View Project

dat-008
Academic Paper Abstract Database
Scientific text (titles, abstracts, authors, publication dates) — Query APIs with keywords to fetch research papers and map publication trends. Stored as JSON / SQLite.
APIPythonResearchText Data
View Details
dat-019
Historical Newspaper OCR Dataset


