vust

Scanned PDF? Read This First

Your Scanned PDF Has No Text. Here's Why AI Can't Summarize It — and the Fix.

A scanned page is a photo of text, not text a computer can read — so @vustSummaryBot (and every text-extraction summarizer) comes back empty on it. This isn't a limit that gets fixed by a bigger model; it needs one OCR step first. Here's how to spot a scanned PDF and the exact workflow to summarize it anyway.

No OCR in SummaryBot · OCR-first workflow belowScope: text extraction only, no OCR
No OCR in SummaryBot4-symptom scanned-PDF checklistOCR-first workflow, step by step

Scope

No AI summarizer reads a scanned PDF directly — SummaryBot included

@vustSummaryBot's PDF handling extracts existing text from the file; on a scanned or image-only PDF there is no text to extract, and the bot returns a clear no-extractable-text message instead of guessing. The same is true of the web /summary/pdf tool and of the PDF-to-Markdown converter. Run OCR first — any tool, see the workflow below — then paste the resulting text in as normal.

This is a genuine file-format limitation, not a quality gap that a future model update closes.
Specimens

See the difference

What happens when you try a scanned PDF directly, how to check for a text layer, and the one-step OCR fix.

Sending a scanned PDF directly

What you'd expect

Upload the scanned PDF to an AI summarizer and get a summary back, same as any other PDF.

What actually happens

The extractor finds zero characters of text — a scanned page is a photograph of text, not encoded text — so the summarizer has nothing to summarize. @vustSummaryBot returns a clear "no extractable text (scanned/image-only)" message instead of guessing or hallucinating a summary from the filename.

Checking for a text layer first

The quick test

Open the PDF and try to select a sentence with your cursor, the way you'd select a sentence in a Word document.

What the result tells you

If a sentence highlights and you can copy-paste it, there's a text layer — any summarizer can read it. If nothing highlights, or the whole page selects as one image, it's scanned/image-only and needs OCR before any AI tool can touch it.

The OCR-first fix

One extra step

Run the scanned pages through an OCR pass to turn the photographed text into real, selectable text.

Then it works normally

Paste the OCR'd text — or a batch of OCR'd pages — into @vustSummaryBot the same way you'd paste any long document, and the normal summary (bullet points, key takeaways, TL;DR) runs exactly as it would on a native-text PDF.
Practical use cases

Who hits the scanned-PDF wall

Researchers with old library scans
A pre-2000 paper only exists as a photocopy-of-a-photocopy PDF from a library archive, with no text layer.
The 4-symptom checklist confirms it's a scan, then the OCR-first workflow gets it into a summarizable state.
Anyone with a photographed document
A contract, form, or report was scanned on a home scanner or photographed on a phone rather than exported digitally.
OCR converts the photographed page into real text before any summarizer — SummaryBot included — can read it.
People who already tried and got an empty result
SummaryBot or the web tool returned a 'no extractable text' message and it wasn't obvious why.
This page explains exactly why that happens and the one extra step that fixes it.
How it works01–03

From scanned PDF to a real summary

  1. 01

    Confirm it's actually scanned

    Try to select a line of text with your cursor — if nothing highlights, it's a scan, not a native PDF.

  2. 02

    Run OCR on the scanned pages

    macOS Preview, Adobe Acrobat, Google Docs, or Tesseract all convert the photographed page into real text; @vustMarkdownBot's photo path works for short scans, one page at a time.

  3. 03

    Paste the extracted text into @vustSummaryBot

    Once you have real text, it is summarized exactly like any native-text PDF — no special mode needed.

Same tool · in Telegram@vustSummaryBot

Once you have text, summarize it here

Run OCR first with any tool above, then paste the result into @vustSummaryBot for the normal bullet/paragraph/key-takeaways/TL;DR summary.

Open in Telegram
Quality & trust

What this page covers

SummaryBot does not OCR PDFs — full stop

On a scanned or image-only PDF the bot tells you it found no text, rather than attempting OCR or guessing at the content. The web /summary/pdf tool has the identical text-extraction-only limitation.

MarkdownBot's OCR path is photo-only, not PDF-only

@vustMarkdownBot can read a Telegram PHOTO (an image file); sending it a scanned PDF document goes through the same text-extraction path as SummaryBot and hits the identical no-text result. Exporting scanned pages as images first is the workaround for short documents.

This is a file-format limitation, not a roadmap gap

No AI summarizer built on text extraction reads a scanned PDF directly — the fix always requires a separate OCR pass first, on any product.

FAQ

Frequently asked questions

Why can't AI read my scanned PDF?

A scanned PDF is a photograph of a page, not encoded text — it's the same file format as a text PDF, but the content is pixels, not characters. Text extraction (what SummaryBot and the web /summary/pdf tool both use) looks for a text layer in the file; if the page is scanned and has no OCR layer already baked in, there's nothing to extract, so there's nothing to summarize. This isn't a bug or a limit that will change with a bigger model — it's a fundamentally different kind of file, and it needs a separate OCR step first.

Does @vustSummaryBot have OCR built in?

No. @vustSummaryBot's PDF handling extracts existing text from a PDF; when a PDF has no extractable text it returns a clear "PDF contains no extractable text (scanned/image-only)" result rather than attempting OCR or guessing. The web /summary/pdf tool works the same way, extracting text client-side with the same no-OCR limitation. If your PDF is a scan, run it through OCR yourself first, then paste the resulting text.

Can any VUST bot OCR my scanned PDF directly?

Not the PDF file itself. @vustMarkdownBot can read text from a PHOTO (an image file), but that input accepts an image rather than a multi-page PDF document. Sending the PDF file itself still needs an extractable text layer. If your scanned document is only a page or two, exporting each page as an image and sending it as a photo is a practical shortcut — see the workflow below.

What's the fastest free way to OCR a scanned PDF before summarizing?

macOS Preview and Adobe Acrobat can both run OCR on a scanned PDF and save a text-searchable version; Google Drive will OCR a PDF on upload if you open it with Google Docs; open-source Tesseract works from the command line for batches. Any of these gets you from "scanned image" to "selectable text" — after that, paste the text into @vustSummaryBot or the /summary/pdf tool exactly as you would any other document.

How do I tell if my PDF is scanned without opening special software?

Try to select a line of text with your cursor the normal way. If it highlights like text in a document, it has a text layer and any AI summarizer can read it directly. If nothing highlights, if a whole page selects as a single block, or if zooming in shows visible paper texture, JPEG artifacts, or a slight rotation on the text, it's a scan — the checklist below covers all four symptoms in detail.

Ready when you are

OCR first, then summarize — same result, one extra step.

A 4-symptom scanned-PDF checklist and the exact OCR-first workflow, so a scan isn't a dead end.