About This Tool
OCR PDF renders scanned pages and uses optical character recognition to produce searchable, copyable English text from image-based documents.
What You Can Do
Scanned-page recognition
Analyze page images when no selectable text exists.
English model
Use the deployed English OCR language model.
Text output
Download recognized text for review or reuse.
How to Use
- 1
Upload a scanned PDF
Upload a scanned PDF.
- 2
Start OCR and allow the English model to load on first use
Start OCR and allow the English model to load on first use.
- 3
Wait while pages are rendered and recognized
Wait while pages are rendered and recognized.
- 4
Review and download the extracted text
Review and download the extracted text.
Practical Use Cases
Scan transcription
Recover typed wording from scanned letters or forms.
Search preparation
Create text that can be indexed alongside an image-only archive.
Privacy and File Processing
OCR PDF keeps its PDF input in browser memory while the operation runs. The workflow creates TXT for local preview or download and does not send the entered values or selected file bytes to the GXA Toolbox application server.
Supported Inputs and Outputs
Input
Output
- TXT
How It Works
PDF.js rasterizes scan pages and Tesseract’s deployed English model recognizes visible characters into downloadable text; it does not add a hidden searchable layer to the source PDF.
Limitations and Important Notes
- Handwriting, low resolution, skew, unusual fonts, and complex layouts reduce accuracy.
- The deployed model recognizes English; other languages are not promised.
- OCR output must be proofread before important use.
Worked Example
Run OCR on a clean 300-dpi English invoice scan, then compare totals and names against the page image.
Helpful Tips
- Use straight, high-contrast scans.
- Never rely on OCR alone for legal, medical, or financial values.
Frequently Asked Questions
Does OCR work on handwriting?
It is designed mainly for printed text; handwriting accuracy may be poor.
Which language is supported?
This deployment enables the English recognition model.
Does OCR change the original PDF?
The workflow extracts text; it does not claim to create a searchable-text-layer PDF.