The PDF has no readable text
A PDF made by scanning paper often contains pictures of words rather than words. Connect tries the text layer first; when there is none, the model reads the page as an image instead. You still get an answer, but it is a reading rather than an extraction — good for what a document says, unreliable for a column of figures you intend to act on without checking.
What the symptom looks like#
Nothing announces itself as broken. The file attaches, the preview shows the pages, and an answer arrives. What gives it away is the quality of the answer: it describes the document accurately in general and goes vague or subtly wrong on specifics — a reference number with a transposed pair of digits, a total that belongs to the row above, a date read off the wrong line.
Searching the PDF in your own viewer is the quickest confirmation. If selecting text is impossible and searching finds nothing, there is no text layer, and everything below applies.
Why it happens#
- The document was scanned
- A scanner produces an image per page and wraps it in a PDF. Unless optical character recognition ran afterwards, no words were ever stored.
- It was photographed and converted
- A phone photo saved as PDF is the same situation with worse geometry.
- The text layer was removed on export
- Some systems flatten a document to images deliberately, usually to stop editing.
- The text is there but not in the pages you asked about
- A mixed document — a typed covering page and scanned attachments — reads well at the front and badly after it.
Connect does not run a separate optical character recognition step and then read the output. The model is shown the page. That is why the result is prose you can question rather than a text file you would have to trust, and also why the failure is quiet rather than loud.
What Connect completed#
- The file was accepted, stored and versioned, with provenance, exactly as any other document is.
- The text layer was attempted first. The image path is a fallback, not the default.
- An answer was produced and cited to the file, so you can put the claim beside the page it came from.
None of that is undone by the missing text layer, and the file remains a perfectly good attachment to a record. What changed is the confidence you should place in any exact value inside it.
What Connect did not complete#
No extraction happened. There is no set of characters lifted out of the document that you or anything else can search, no table pulled into rows and columns, and nothing that will feed an import. A scanned price list cannot become records on a sheet by attaching it, because the values were never read as values.
Connect also does not silently improve the file. It is not re-scanned, sharpened or re-saved with a text layer added, so a second question about the same PDF meets the same conditions as the first.
What to do about it#
Ask narrower questions — one field at a time rather than a summary of the document.
Result A single value is far easier to check, and a model asked for one thing tends to look in one place.
Open the file in preview beside the answer before acting on any number.
Result You compare the claim against the source rather than against your memory of it.
Go back for the original wherever one exists — the spreadsheet, the invoice from the accounting system, the export rather than the print-out.
Result Reading tables out of a file works on a real table and produces rows rather than sentences.
If paper is genuinely the only source, run optical character recognition in your scanner software and attach the searchable PDF it produces.
Result The text layer exists from then on, and every later question about the document takes the accurate path.
An administrator has no setting to change here — there is no workspace option that adds a text layer or switches the fallback off. What an administrator can usefully do is fix it upstream: a scanner configured to save searchable PDFs removes the problem for everyone who uses it.
Escalate only when a PDF that genuinely has selectable text is still being read as an image, which shows up as vague answers on a document you can search in your own viewer. That is a different problem from this one and worth reporting with the format and the page count.
Questions#
Will Connect tell me it fell back to reading the page as an image?
Do not rely on it saying so. The safer habit is the reverse: when a figure from a PDF matters, check whether the PDF is searchable before you trust the number. Connect saying it cannot make out a value is a correct answer and a better outcome than a confident guess at a blurred one.
Is a photo of the page better or worse than the scanned PDF?
About the same, and often worse, because a phone photo adds perspective and shadow to the problems the scan already has. Neither is close to the original document. If you have a choice, send the file the document came from.
Can I attach the scan and the spreadsheet together?
Yes, and it is a sensible pattern when the scan is the authority and the spreadsheet is the workable copy. Ask your question against the spreadsheet and keep the scan attached to the record as evidence of where the numbers came from.