06/26/2026
Nine structured HuggingFace datasets are in production — including 151,120 canonical person records, 19,703 place records with geocoordinates, and a co-occurrence embedding for every person named in the War of the Rebellion Official Records. The last systematic finding aid for this primary source was the 1901 General Index. Here’s what it takes to build the next one.
The War of the Rebellion OR (1880–1901) is the largest primary-source compilation ever produced on the American Civil War: 128 volumes, 130 serials, approximately 138,000 pages of Union and Confederate correspondence, battle reports, orders of battle, strength and casualty returns. Researchers have cited it for 125 years. Few have searched it at scale.
The first problem is OCR. These are scans of 19th-century government prints — quality varies across 128 books from 21 years of printing on different equipment. We built a multi-engine arbitration pipeline (Apple Vision, Tesseract, MLX Falcon, Surya, olmOCR) running against a quality scorer that accepts no candidate without a measured improvement margin. Recent GPU smoke test: RTX 4090, 100 representative pages, $0.33 total. olmOCR improved 78 of those pages; Surya improved 77; combined two-engine best: 85. We’ve scored 139,362 retained run-of-text pages and flagged ~40,000 for targeted remediation. Stage-01 triage logged 79,636 candidate cleanup rows and 117,372 individual OCR-hit occurrences — sorted into deterministic-fix (141 rows), noise-suppress (17,156), and model-triage (62,339) lanes.
The second problem is authority control. The 1901 General Index is the only systematic finding aid ever produced for the OR — and it systematically omitted USCT and African American content. We’re building: 151,120 canonical person records with variant name normalization, 19,703 place records with geocoordinates and theater assignment, 6,541 unit records, 2,156 dated events with OR serial:page locators, and 20,306 period vocabulary terms — plus a dedicated USCT cross-reference that closes the gap the 1901 index left open.
Output: 174 physical volumes through LSI distribution, 15+ reference apparatus volumes (9 Tier 1 structural finding aids, 6+ Tier 2 thematic enrichments), and 9 HF datasets under CC BY 4.0 — including the first person co-occurrence embedding corpus for 19th-century American military history.
Limitation: the authority tables are machine-extracted and require human spot-check before serving as scholarly ground truth. Every OCR candidate clears a quality gate before entering production. This is version 1 of infrastructure that will require continued iteration.