Scanned PDF? Read This First
Your Scanned PDF Has No Text. Here's Why AI Can't Summarize It — and the Fix.
A scanned page is a photo of text, not text a computer can read — so @vustSummaryBot (and every text-extraction summarizer) comes back empty on it. This isn't a limit that gets fixed by a bigger model; it needs one OCR step first. Here's how to spot a scanned PDF and the exact workflow to summarize it anyway.
Scope
No AI summarizer reads a scanned PDF directly — SummaryBot included
@vustSummaryBot's PDF handling extracts existing text from the file; on a scanned or image-only PDF there is no text to extract, and the bot returns a clear no-extractable-text message instead of guessing. The same is true of the web /summary/pdf tool and of the PDF-to-Markdown converter. Run OCR first — any tool, see the workflow below — then paste the resulting text in as normal.
See the difference
What happens when you try a scanned PDF directly, how to check for a text layer, and the one-step OCR fix.
Who hits the scanned-PDF wall
- Researchers with old library scans
- A pre-2000 paper only exists as a photocopy-of-a-photocopy PDF from a library archive, with no text layer.
- The 4-symptom checklist confirms it's a scan, then the OCR-first workflow gets it into a summarizable state.
- Anyone with a photographed document
- A contract, form, or report was scanned on a home scanner or photographed on a phone rather than exported digitally.
- OCR converts the photographed page into real text before any summarizer — SummaryBot included — can read it.
- People who already tried and got an empty result
- SummaryBot or the web tool returned a 'no extractable text' message and it wasn't obvious why.
- This page explains exactly why that happens and the one extra step that fixes it.
From scanned PDF to a real summary
- 01
Confirm it's actually scanned
Try to select a line of text with your cursor — if nothing highlights, it's a scan, not a native PDF.
- 02
Run OCR on the scanned pages
macOS Preview, Adobe Acrobat, Google Docs, or Tesseract all convert the photographed page into real text; @vustMarkdownBot's photo path works for short scans, one page at a time.
- 03
Paste the extracted text into @vustSummaryBot
Once you have real text, it is summarized exactly like any native-text PDF — no special mode needed.
Once you have text, summarize it here
Run OCR first with any tool above, then paste the result into @vustSummaryBot for the normal bullet/paragraph/key-takeaways/TL;DR summary.
What this page covers
SummaryBot does not OCR PDFs — full stop
On a scanned or image-only PDF the bot tells you it found no text, rather than attempting OCR or guessing at the content. The web /summary/pdf tool has the identical text-extraction-only limitation.
MarkdownBot's OCR path is photo-only, not PDF-only
@vustMarkdownBot can read a Telegram PHOTO (an image file); sending it a scanned PDF document goes through the same text-extraction path as SummaryBot and hits the identical no-text result. Exporting scanned pages as images first is the workaround for short documents.
This is a file-format limitation, not a roadmap gap
No AI summarizer built on text extraction reads a scanned PDF directly — the fix always requires a separate OCR pass first, on any product.
Frequently asked questions
Why can't AI read my scanned PDF?
A scanned PDF is a photograph of a page, not encoded text — it's the same file format as a text PDF, but the content is pixels, not characters. Text extraction (what SummaryBot and the web /summary/pdf tool both use) looks for a text layer in the file; if the page is scanned and has no OCR layer already baked in, there's nothing to extract, so there's nothing to summarize. This isn't a bug or a limit that will change with a bigger model — it's a fundamentally different kind of file, and it needs a separate OCR step first.
Does @vustSummaryBot have OCR built in?
No. @vustSummaryBot's PDF handling extracts existing text from a PDF; when a PDF has no extractable text it returns a clear "PDF contains no extractable text (scanned/image-only)" result rather than attempting OCR or guessing. The web /summary/pdf tool works the same way, extracting text client-side with the same no-OCR limitation. If your PDF is a scan, run it through OCR yourself first, then paste the resulting text.
Can any VUST bot OCR my scanned PDF directly?
Not the PDF file itself. @vustMarkdownBot can read text from a PHOTO (an image file), but that input accepts an image rather than a multi-page PDF document. Sending the PDF file itself still needs an extractable text layer. If your scanned document is only a page or two, exporting each page as an image and sending it as a photo is a practical shortcut — see the workflow below.
What's the fastest free way to OCR a scanned PDF before summarizing?
macOS Preview and Adobe Acrobat can both run OCR on a scanned PDF and save a text-searchable version; Google Drive will OCR a PDF on upload if you open it with Google Docs; open-source Tesseract works from the command line for batches. Any of these gets you from "scanned image" to "selectable text" — after that, paste the text into @vustSummaryBot or the /summary/pdf tool exactly as you would any other document.
How do I tell if my PDF is scanned without opening special software?
Try to select a line of text with your cursor the normal way. If it highlights like text in a document, it has a text layer and any AI summarizer can read it directly. If nothing highlights, if a whole page selects as a single block, or if zooming in shows visible paper texture, JPEG artifacts, or a slight rotation on the text, it's a scan — the checklist below covers all four symptoms in detail.
Related document tools.
Ready when you are
OCR first, then summarize — same result, one extra step.
A 4-symptom scanned-PDF checklist and the exact OCR-first workflow, so a scan isn't a dead end.