PDF is the universal container: that convenience is exactly why it causes trouble. Under one extension you may find an electronic document, a scan, a permission-protected file, or nothing but images. Most PDF questions land in one of four buckets: it will not open, search finds nothing, it is too large, or merging wrecked the structure.
The first question in every case is whether the document is electronic or scanned, because it determines what is possible downstream. An electronic document has a real text layer that can be copied, searched and edited. A scan is an image with an invisible OCR layer attached, and that layer has three properties: accuracy typically runs 85–98%, errors concentrate in rare characters, punctuation and digits, and the text is not editable — what you copy out is the OCR's guess, not the source. To tell them apart, try selecting and copying a passage: if nothing sensible comes out, it is a scan. Any workflow that needs to edit content must OCR first, and OCR quality caps everything downstream.
Encryption is routinely misunderstood as a single switch when it is actually two layers. The owner password decides whether the file can be opened at all; the user password only restricts what you may do once it is open. So "will not open" and "opens but I cannot copy" are different problems requiring different fixes. Removing an owner password requires the original one — that is a property of the format, not a tool limitation, which is why decryption tooling applies to files where the password is known.
Search on scans has a subtle signature: it appears to work, and the results are untrustworthy. Searching lO1 misses what should match 101, because OCR read a digit as a letter. If the goal is archival retrieval, image-level search that matches whole pages visually beats relying on the OCR layer. If the goal is content editing, OCR first and then hand-check the fields where errors hurt most: amounts, identifiers and dates.
Size has an extreme composition: about 95% of a scanned PDF is image, with the text layer measured in tens of kilobytes. This explains why generic compress-document passes barely help — they optimise fonts, objects and metadata, which were never the bulk. The real lever is the embedded image: drop 300dpi to 150dpi, replace lossless PNG with lossy JPEG, and convert single-colour scans to greyscale or bitonal. Together these typically save 60–80%. A PDF compressor working at the image level is what you want, and note that some tools default to structural compression, so choose the image option explicitly.
Merging is where structure gets sacrificed. Bookmarks, section jumps, form field values and internal links all depend on page and object references that a naive merge reorders. Bookmarks can end up pointing at an unrelated but existing page, links break, and form values disappear. Use a tool that explicitly preserves the outline, then verify each bookmark and link afterwards. If the source documents cross-reference each other, do not merge at all — serve separate PDFs behind an index page, which is far more reliable than repairing links. Splitting carries the same risk for the same reason, since page renumbering breaks the same references.