"Are our documents ready for AI?" is the first question a manufacturer asks — and the only wrong answer is an opinion. Readiness is measured. Here are the five dimensions we evaluate on every corpus, with the thresholds that separate a reliable assistant from one that invents.
A "clean" PDF yields its text directly; a scan from 1998 has to go through visual reading (OCR), with a measurable confidence rate. The classic trap: PDFs that seem to have text, but whose text layer from an old OCR is riddled with errors — they silently pollute every answer.
An assistant cites "Manual V3 › Troubleshooting › page 42" only if your documents have detectable sections. Blocks of text without headings, flattened tables or presentations without hierarchy yield "floating" excerpts that are impossible to cite precisely.
The "final-v2-FINAL.pdf" copy next to the original is not harmless: two versions of the same content diverging by a paragraph give contradictory answers depending on which version the assistant pulls. They must be detected beforehand, then you decide which one governs — not discover them in a bad answer.
A troubleshooting corpus needs procedures and safety instructions. If it only contains sales sheets and parts catalogs, the assistant will be excellent… at answering questions no one asks. Coverage is verified by content type, not by the weight of the corpus.
The most overlooked dimension: what is missing. The breakdowns your technicians resolve out of habit, passed on by word of mouth, exist in no document. A good system reveals them — every unanswered question is logged and becomes a draft to be validated by your experts. The blind spot gets filled instead of repeating itself.
That's exactly what our readiness audit does: a score out of 100, the state of each document and a prioritized action plan — the report is yours, whether you continue with us or not.
Request an audit of my documents