← Back to Blog

Working with PDFs: The Two Passwords, OCR on Scans, and Bookmarks Lost in Merging

PDF is the universal container: that convenience is exactly why it causes trouble. Under one extension you may find an electronic document, a scan, a permission-protected file, or nothing but images. Most PDF questions land in one of four buckets: it will not open, search finds nothing, it is too large, or merging wrecked the structure.

The first question in every case is whether the document is electronic or scanned, because it determines what is possible downstream. An electronic document has a real text layer that can be copied, searched and edited. A scan is an image with an invisible OCR layer attached, and that layer has three properties: accuracy typically runs 85–98%, errors concentrate in rare characters, punctuation and digits, and the text is not editable — what you copy out is the OCR's guess, not the source. To tell them apart, try selecting and copying a passage: if nothing sensible comes out, it is a scan. Any workflow that needs to edit content must OCR first, and OCR quality caps everything downstream.

Encryption is routinely misunderstood as a single switch when it is actually two layers. The owner password decides whether the file can be opened at all; the user password only restricts what you may do once it is open. So "will not open" and "opens but I cannot copy" are different problems requiring different fixes. Removing an owner password requires the original one — that is a property of the format, not a tool limitation, which is why decryption tooling applies to files where the password is known.

Search on scans has a subtle signature: it appears to work, and the results are untrustworthy. Searching lO1 misses what should match 101, because OCR read a digit as a letter. If the goal is archival retrieval, image-level search that matches whole pages visually beats relying on the OCR layer. If the goal is content editing, OCR first and then hand-check the fields where errors hurt most: amounts, identifiers and dates.

Size has an extreme composition: about 95% of a scanned PDF is image, with the text layer measured in tens of kilobytes. This explains why generic compress-document passes barely help — they optimise fonts, objects and metadata, which were never the bulk. The real lever is the embedded image: drop 300dpi to 150dpi, replace lossless PNG with lossy JPEG, and convert single-colour scans to greyscale or bitonal. Together these typically save 60–80%. A PDF compressor working at the image level is what you want, and note that some tools default to structural compression, so choose the image option explicitly.

Merging is where structure gets sacrificed. Bookmarks, section jumps, form field values and internal links all depend on page and object references that a naive merge reorders. Bookmarks can end up pointing at an unrelated but existing page, links break, and form values disappear. Use a tool that explicitly preserves the outline, then verify each bookmark and link afterwards. If the source documents cross-reference each other, do not merge at all — serve separate PDFs behind an index page, which is far more reliable than repairing links. Splitting carries the same risk for the same reason, since page renumbering breaks the same references.

Advertisement

Frequently Asked Questions

What is the difference between the owner password and the user password in PDF encryption?

**They control entirely different things.** The owner password decides whether the file can be opened at all — without it you cannot get in. The user password only decides what you may do after opening: the file opens but printing or copying is blocked. The common complaint of a document that will not open is the former; opens-but-cannot-copy is the latter, and **you must establish which one you are facing before debugging**. A PDF encryption tool lets you set both layers separately, and a PDF decrypt tool handles already-protected files — but removing an owner password requires the original one, which is a design limit rather than a tool defect.

Can scanned PDFs be searched reliably?

**Yes, but the results cannot be trusted.** A scan is an image; the searchable text layer is OCR output at typically 85–98% accuracy, with errors concentrated in rare characters, punctuation and digits. The consequences are concrete: searching for lO1 (lowercase L, capital O, digit one) misses what should match 101, and copying OCR text into a data pipeline yields corrupted identifiers. **For archival search only**, use image-level search that matches whole pages visually rather than relying on the text layer. **For content editing**, OCR first and manually spot-check the critical fields — amounts, identifiers and dates, where errors hurt most.

How do I compress a 20MB PDF down to 5MB?

**Compress the images, not the document structure.** In a scan roughly 95% of the bytes are images and the text layer is only tens of kilobytes, so a generic compress-document pass over fonts, objects and metadata changes almost nothing. What works is recompressing the embedded images: drop 300dpi to 150dpi, or replace lossless PNG with lossy JPEG, saving 60–80% with almost no visible difference. A PDF compressor tool operates at exactly that level. Note that some tools default to structural compression, so you may need to pick the image-compression option explicitly.

After merging several PDFs, the table of contents and internal links all break. Can that be fixed?

**Broken links can be fixed; misaligned bookmarks are much harder.** Simple merge tools discard the original outline and in-document links, so a bookmark may end up pointing at the wrong page. A tool that preserves the outline keeps most of it, and you should still verify each entry after merging. If the source documents cross-reference each other, prefer not to merge at all — serve separate PDFs behind an index page, which is more reliable than repairing links. A practical extra check is to export the merged text with a PDF extraction tool and confirm the section order matches expectations.

← Back to Blog