Text PDF vs Scanned PDF: Why the Difference Matters for BOQs
Last reviewed:
Why Document Quality Matters
A text PDF exposes embedded characters for supported capture, but its tables and values still require review. A scanned PDF contains page images and needs OCR or manual transcription before its text can be processed.
Identifying the document type immediately dictates the necessary quality control workflow.
Text PDFs vs Scanned PDFs
- Native Text PDF: Generated directly from software (like Word or Excel) via "Save as PDF". You can usually select individual letters and numbers, allowing software to access the embedded text while still requiring structural and value checks.
- Scanned Image PDF: Created when a physical piece of paper is run through a scanner. It is essentially a photograph. You cannot highlight the text. Software must use Optical Character Recognition (OCR) to "guess" what the shapes mean.
Comparison Overview
| Feature | Native Text PDF | Scanned Image PDF |
|---|---|---|
| Selectable Text | Yes | No |
| Extraction Accuracy | Varies by layout; review required | Variable (depends on scan quality) |
| Processing Method | Direct character extraction | Requires OCR technology |
| File Size | Varies with embedded content | Varies with resolution and compression |
| Review Requirement | Standard structural review | Rigorous line-by-line verification |
A Practical Example
A hypothetical consultant prints a BOQ, signs it with a pen, and scans it back into the computer. This is now a scanned PDF.
If the page is skewed or the resolution is low, OCR software might misread an "8" as a "3". A digitally exported PDF would retain an embedded text layer, although its table structure and values would still require review.
Limitations of Scanned Documents
Scans often suffer from skew, compression artifacts and handwriting over text. These elements can reduce recognition confidence.
No software can guarantee perfect extraction from a poor-quality scan. Human validation is always required.
How Quantara Processes PDFs
Quantara accepts text-based and scanned PDFs. It stores extractable text from text-based PDFs and creates review candidates only from supported detected table rows; plain paragraph text does not become BOQ candidates. Scanned or image-only documents are detected and flagged as requiring OCR, but OCR text extraction is not currently available, so scanned content requires manual transcription.
Quantara currently focuses on supported document extraction, BOQ structuring, project organization, templates, revisions, and professional outputs.
Professional Disclaimer
This information is provided for general educational purposes and does not replace project-specific advice or professional judgment. Quantities, units, specifications, rates, assumptions, exclusions and project documents must be reviewed by an appropriately qualified construction professional before tender, procurement, contractual or construction use.
Frequently Asked Questions
How do I know if my PDF is text or scanned?
Try to highlight a single word with your mouse cursor. If you can select individual letters, it is a text PDF. If clicking highlights the entire page as a block, it is a scanned image.
Can a PDF be a mix of both?
Yes. Some documents contain digitally generated text pages alongside appended scanned appendices or stamped signature pages.
What is OCR?
OCR stands for Optical Character Recognition. It is technology that analyzes the shapes in an image and translates them into machine-readable text.
Why do scans result in larger file sizes?
A scan stores page imagery, while a text PDF can store characters and vector instructions. Actual file size varies with images, fonts, resolution, compression and other embedded content.
Can I request a text PDF from the client?
Yes. It is standard practice during tender periods for contractors to request the native files or text-based PDFs to reduce administrative burden and pricing errors.
Related Reading
Explore Related BOQ Workflows
Quantara helps construction teams turn supported project documents into structured BOQ records, controlled templates, revisions and professional outputs.