Can a local AI read scanned PDFs?
Scans are images, not text, which is exactly why so many archives sit unread on disks. A local tool can read them without sending a single page anywhere.
Yes. A scanned PDF is a picture of a page, so a local tool reads it with OCR running on your machine, or with a vision model that looks at the page directly. Both work entirely offline: the scan is never uploaded to be read. Once the text is extracted, the scan behaves like any other document in the library, searchable and quotable, with citations pointing back to the original page.
Why scans break ordinary tools.
A born-digital PDF carries its text inside it. A scanned PDF carries only a photograph of the page, and a tool that searches photographs finds nothing. That is why a decade of archive boxes, court filings, old journal articles and handwritten field notes sit in folders that no search box can reach, and why the usual cloud answer, upload them to a service that OCRs them, is a non-starter for anything confidential.
Reading a scan locally takes one of two forms. OCR, optical character recognition, turns the image into text; Windows ships an OCR engine, so a local tool can use it without any download or account. A vision model does it the other way: it looks at the page image directly and reads what it sees, which also covers layouts that trip plain OCR.
What works, and what does not.
Clean, printed pages are the easy case, and modern OCR handles them well: the extracted text is accurate enough to search, retrieve, and quote. The honest limits are physical. Skewed scans, faint photocopies, two-column layouts, tables, marginalia and handwriting all reduce accuracy, and handwriting most of all. A scan is only as readable as its photograph.
The working habit is proportionate to the risk: trust the extracted text for search and retrieval, and spot-check it against the page image before quoting anything important. In a grounded tool the citation makes that check a click, because the answer points at the passage in the scan itself.
How scans join the loop.
Once extracted, a scan stops being a special case. Its text is chunked and embedded like any other document, so questions retrieve its passages alongside everything else in the library, and the citation leads back to the scanned page, not to a degraded copy of it. Scans are one entry in the full list of formats a local tool reads, and the pattern is the same for all of them. A field notebook photographed at the site, an archive scan, and a downloaded paper end up in one index, answered together, on a machine that never uploaded any of them.