Skip to content
Knowledge base10 min read

Scanned documents and images

Photos and scanned PDFs carry no text, so Agentency transcribes them before training. Arabic scans included.

Browse topics

What it is

Some documents contain no text at all. A photograph of a menu, a page scanned on an office copier, a PDF that is really just a picture of paper — to a computer these are pictures, not words. Copy-and-paste from them gives you nothing.

Agentency reads those documents anyway. When you upload a file that has no usable text inside it, the file is sent through text recognition (often called OCR, for optical character recognition) before anything else happens. Recognition transcribes what is printed on each page into ordinary text, and only then does the normal training run. From that point on the document behaves like any other source: it is searched, quoted, and cited exactly like a Word file you typed yourself.

You do not switch this on. You do not pick a language. You upload the file the same way you upload anything else, and Agentency decides whether the document needs transcribing.

When you would use it

Use this whenever the only copy you have is a picture of the words:

  • A photo or screenshot of a price list, a certificate, a leaflet, a printed FAQ card.
  • A scanned PDF — a contract, a policy, an old brochure that was fed through a scanner years ago.
  • An Arabic PDF. Arabic PDFs are a special case worth knowing about: many of them technically carry a text layer, but that layer is scrambled — letters in the wrong order, words reversed, characters that look right on screen and come out as gibberish when copied. Agentency does not trust an Arabic text layer for exactly this reason. It transcribes the page from the image instead, which produces clean, correctly-ordered Arabic.

If your document already has real, selectable text, none of this applies. Open it, try to select a sentence, and press copy. If you get the sentence, the file has a text layer and will be read directly — see Upload files.

Where to find it

Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Files.

There is no separate screen. Recognition happens inside the normal upload flow, and you watch it on the Knowledge table.

Which file types get transcribed

Images, always. The upload picker accepts .png, .jpg, .jpeg, and .webp. An image can never have a text layer, so every accepted image goes through recognition.

PDFs, only when they need it. Agentency inspects the PDF first and transcribes it when any of these is true:

  • there is no text layer at all (a pure scan);
  • there is barely any text — a handful of characters across a whole page, which is what a scan with a stray header looks like;
  • the text is Arabic-dominant, for the ordering reason above.

Everything else — a normal PDF exported from Word, a text file, a spreadsheet, a .docx, an .odt — is read straight from its own contents and never goes near recognition.

Steps

  1. Open Dashboard → Chatbots → your chatbot → Knowledge → Add → Files.
  2. Give the source a clear title. This title is what appears in answers when the chatbot cites the document, so "2026 Price List" beats "scan_0042".
  3. Choose the image or PDF from your device.
  4. Leave Start training on and upload.
  5. Watch the row on the Knowledge table. It will show Reading Document while the transcription runs.
  6. When that clears, the row moves into normal training and finally reads as trained.
  7. Ask Test Chatbot a question using a phrase that is actually printed on the page. If it answers, the transcription worked.

What you will see

A file that needs recognition shows a distinct badge — Reading Document — with a moving progress bar that has no percentage. The tooltip explains it: "Transcribing this scanned or image-based document into text before training — this can take a moment."

The bar has no percentage on purpose. Recognition is a single job that either finishes or does not; there is no honest halfway number to show you, so Agentency shows movement rather than a fake figure.

After Reading Document clears, the row continues through the ordinary training states — pending, then training, then trained. See Dataset status and training for what each badge means.

Why it takes longer

A normal upload is read while you wait: you press upload, the text comes out of the file in the same moment, and training begins. A scan cannot work that way. Every page has to be examined and transcribed, and that work is queued and runs in the background on its own, separate from the rest of the site.

Practically, this means three things:

  • A scanned document is slower than the same document as real text. Minutes rather than seconds is normal, and a long document is longer still.
  • You can close the tab. The work does not depend on your browser staying open. Come back later and read the badge.
  • Nothing else is blocked. You can keep adding other sources while a scan is being read. Agentency will not run two jobs on the same document at once, so a scan and its training never trip over each other.

If a transcription is interrupted — a restart, a network problem at the recognition service — it is retried automatically with a delay between attempts, and a retry never re-transcribes a document that already succeeded. A job that somehow got lost entirely is picked up and re-run by a background check after about 45 minutes, so a row sitting on Reading Document for an hour is unusual but not permanently stuck.

How to tell whether it worked

Three checks, in order:

  1. The badge. Trained means text came out and was used. Failed means it did not.
  2. Ask a real question. Open Test Chatbot and ask something whose answer only exists on that page. Use a phrase printed on the document, not a paraphrase.
  3. Check the citation. If the answer cites your document by the title you gave it, the transcription is genuinely in the chatbot's knowledge.

If the row failed, open it. The error is written in plain language. The one you will see most often for a scan is "No trainable content was found in this item." That is the honest verdict: Agentency looked at every page and found nothing readable.

When a scan extracts nothing useful

"No trainable content was found in this item." means one of these:

  • The page is genuinely blank, or is a photograph with no writing in it (a logo, a product shot, a signature).
  • The photo is too blurred, too dark, too small, or too skewed for the letters to be made out.
  • The writing is handwritten. Handwriting is not reliable.
  • The text is a decorative or heavily stylised typeface.

What to do, in the order that costs you least:

  1. Re-take the picture properly. Flat on a table, straight on rather than at an angle, no shadow across the page, and fill the frame with the page. A phone camera in good light beats a bad scanner.
  2. Rescan at a higher quality if you have the original on paper. Scan as PDF at a normal document setting rather than a "small file" setting.
  3. Split a big document. One page that fails is easier to diagnose than a fifty-page file that fails as a whole. Upload one page, confirm it reads, then do the rest.
  4. Give up on the image and type it. For a short page — opening hours, a price table, a return policy — pasting the text into Add text knowledge is faster than fighting a bad scan and gives a perfect result.
  5. Ask the author for the original. If the scan came from a Word document or an InDesign export, the original file will always produce better knowledge than a picture of the printout.

If you fixed the image, use Replace file on the existing row rather than deleting and re-adding. Replacing keeps the title and the collection, and re-runs recognition on the new file.

The honest limitation

Recognition transcribes what it can see. It cannot see what is not there.

A crisp scan of a clean printed page gives you knowledge that is as good as if you had typed it. A hurried phone photo of a crumpled page, taken at an angle in poor light, gives you knowledge full of missing words and misread characters — and the chatbot will answer confidently from that mangled text, because it has no way of knowing the text is wrong. This is the failure mode that actually hurts, because it is quiet: nothing shows as failed, and the answers are simply subtly incorrect.

So treat a photograph of text as the last resort, not the first. Order of preference, best to worst:

  1. The original digital file (.docx, a text PDF, a spreadsheet).
  2. Text pasted directly into the chatbot.
  3. A clean, flat, well-lit scan.
  4. A phone photograph.

And whichever you use, always test with a real question afterwards. A scan you never tested is a scan you are only assuming works.

Limits and plan notes

  • Recognition does not change the upload ceilings. Each file is still capped at 10 MB, a batch at 20 files, and the whole batch at 90 MB in total.
  • Transcribed text counts toward your chatbot's knowledge cap exactly like typed text — Free 400 KB, Starter 10 MB, Standard 20 MB, Pro 40 MB, Agency 60 MB. A dense scanned manual can be much larger than the file size suggests, because a 2 MB image of a page can produce several thousand characters of text. See Knowledge limits.
  • The size check for a scan happens after transcription, because the amount of text is unknown until the page has been read. A scan can therefore be rejected for exceeding your plan's knowledge cap only at the end. That is not a bug; it is the first moment the real size is known.
  • Recognition is a platform capability, not a plan feature. If it has been turned off for a deploy, images are rejected at upload instead of being transcribed, and PDFs fall back to whatever text layer they have.

Common problems

The row says Reading Document and has for a long time.

Give a long document time — this is background work and a large scan is genuinely slow. If it is still there after an hour, an automatic recovery pass will re-run it.

The row failed and mentions the training service.

The recognition service was temporarily unreachable. This retries by itself; if the row is still failed after a while, retry it from the row.

The Arabic came out in the wrong order.

That is the problem recognition exists to solve, so it should not happen on a scan. If it does, the PDF probably had just enough Latin text to be read directly instead of transcribed. Save the page as an image and upload that, which forces recognition.

The answers quote the document but get numbers wrong.

Misread characters in a poor scan. Numbers and tables suffer worst — a 5 read as 6 in a price list is invisible until a customer complains. Check the transcribed figures against the original, and for anything price-critical, paste the numbers in as text instead.

I uploaded a photo and it was rejected at upload.

Either the format is not one of .png, .jpg, .jpeg, .webp, or the file is over 10 MB, or image reading is off for this deploy. Convert or shrink the image, or use Add text knowledge.

Common questions

Which files get transcribed instead of read directly?

Every image (.png, .jpg, .jpeg, .webp), and any PDF with no usable text layer — a pure scan, a near-empty text layer, or an Arabic-dominant document.

Why are Arabic PDFs transcribed even when they contain text?

Arabic text layers are frequently scrambled — letters reordered, words reversed — so copying them produces gibberish. Reading the page as an image gives clean Arabic.

How do I know whether the transcription worked?

The row reaches trained rather than failed, and Test Chatbot can answer a question using a phrase printed on the page and cite your document by title.

My scan failed with 'No trainable content was found'. What now?

Nothing readable was found. Re-take the photo flat and well lit, rescan at a higher quality, or type the content in as text knowledge instead.

Is a phone photo of a page good enough?

It is the worst of the options. A blurred or angled photo produces missing and misread words, and the chatbot answers confidently from that mangled text.

Was this article helpful?

Ready to try it on your own content?

Create a free workspace, add a document, and ask the questions your team is tired of answering.

Scanned documents and images | Agentency Help