Recurring text-quality defects across six languages, measured as events per 10,000 extracted characters to show where non-Latin script handling falls furthest behind.
The dashboard compares recurring extraction problems that persist after PDF text has already been pulled into plain text, focusing on the rate of cleanup-worthy artifacts rather than on document counts.
| Language | Script | Noise | Mid-word | Spaces | OCR | Total / 10k | Total chars |
|---|