Historical PDF Extraction Failures

Recurring text-quality defects across six languages, measured as events per 10,000 extracted characters to show where non-Latin script handling falls furthest behind.

Highest total failure rate
Non-Latin average
Dominant defect class
Corpus coverage
Total recurring failure load by language
Stacked rates reveal how the overall burden and the defect mix change across scripts. Languages are sorted by total failures per 10k characters.
Failure-rate multiples relative to English
Each cell shows how much a failure type rises or falls versus the English rate. The visual gap widens sharply on Russian, Arabic, and Chinese.
Average defect profile: Latin vs non-Latin scripts
A family-level view shows the quality gap is systematic, not just a single-language outlier.
Key readout

The dashboard compares recurring extraction problems that persist after PDF text has already been pulled into plain text, focusing on the rate of cleanup-worthy artifacts rather than on document counts.

Highest noise rate
Highest mid-word break rate
Highest excessive-space rate
Highest OCR-artifact rate
Per-language rate table
Exact rates and corpus size for each language sample.
Language Script Noise Mid-word Spaces OCR Total / 10k Total chars