Processing PDFs with a vision LLM means treating each page as an image and passing it to a multimodal model rather than extracting the text layer first. This approach handles PDFs that break standard text extraction: scanned documents with no text layer, PDFs with complex layouts where text extraction produces garbled output, files with important content in figures and tables, and mixed documents where visual and textual content need to be understood together. The pipeline is more involved than text extraction but gives you a fundamentally more capable document processing system.
When to Use Vision Over Text Extraction
Standard PDF text extraction (pdfplumber, pypdf, pdfminer) works well for digitally-created PDFs with a clean text layer and simple layout. It fails or degrades on four common document types that are better handled with vision. Scanned PDFs: documents photographed or scanned to PDF have no embedded text — the pages are images and text extraction returns nothing. Complex multi-column layouts: academic papers, newsletters, and formatted reports with multi-column text are often extracted in wrong reading order by text tools, mixing content from different columns. Tables with merged cells: complex tables lose their structure during text extraction, making the extracted content unusable for analysis. Figure-heavy technical documents: engineering specs, scientific papers, and financial reports often have more information in figures than in prose — text extraction misses all of it. The simplest diagnostic: run pdfplumber on a document and check whether the extracted text is coherent and complete. If it is, use text extraction. If it is garbled, missing, or incomplete, use vision processing.
The Page-as-Image Pipeline
The core pipeline converts each PDF page to an image and processes it through a vision LLM. PyMuPDF (fitz) handles the conversion: open the PDF, iterate over pages, render each as a PNG image at the desired resolution, and save or pass directly to the model. Resolution matters: 150 DPI balances quality and file size for most documents; 200–300 DPI is better for small-text documents or when OCR accuracy is critical. The rendered image is then base64-encoded and passed to the vision model API alongside an extraction prompt. For a 20-page document, this produces 20 separate API calls — which can be parallelised for throughput using asyncio or a thread pool.
Extraction Libraries That Do This For You
Several libraries abstract the page-rendering-plus-vision-model pipeline. Docling (IBM, open source) handles PDF ingestion with layout analysis, table extraction, and figure detection, outputting structured markdown or JSON with position metadata. It uses vision models internally for complex layout elements and rule-based extraction for clean text — a hybrid that optimises cost without sacrificing quality on difficult content. Marker converts PDFs to markdown with high fidelity, handling multi-column layouts and tables correctly, and uses Surya (a layout detection model) to understand page structure before extracting text. LlamaParse (LlamaIndex’s commercial service) uses GPT-4V to parse complex PDFs into structured output, particularly strong on financial documents and research papers. For teams that do not want to build and maintain a custom pipeline, these libraries provide production-quality PDF processing with minimal code.
Figure 1 — PDF processing approaches: complexity vs capability
Structuring the Extraction Output
Raw page-by-page text extraction is often not what you need downstream — you need structured output: headings, sections, tables as data, figures as captioned items. Prompting the vision model with structure requirements produces more usable output: “Process this PDF page and return a JSON object with these fields: page_type (cover/content/appendix), headings (list of heading text), body_text (main paragraph content), tables (list of tables, each as a list of rows), figures (list of figure captions or descriptions).” This structured extraction integrates naturally with downstream processing — the tables field can be inserted directly into a database, the headings build a document outline, and the figures list feeds a multimodal RAG index.
Cost Management at Scale
Processing large PDF collections with a vision LLM requires cost discipline. GPT-4o processes approximately 1,000 tokens of visual input per page at standard resolution — at $2.50/M input tokens, that is $0.0025 per page for image tokens plus text prompt and output tokens, totalling roughly $0.005–0.015 per page depending on output length. For a 100,000-page archive, cloud API costs reach $500–1,500. Cost optimisation strategies: use a cheaper model (Claude Haiku, Gemini Flash) for a first pass to filter out blank pages, cover pages, and pure image pages with no text; use text extraction first and route only the pages that produce garbled or empty text to the vision pipeline; run the vision model locally (InternVL2-8B on a single 24GB GPU) for high-volume processing where hardware is available. The hybrid approach — cheap text extraction for clean PDFs, vision processing for difficult ones — typically reduces vision processing volume by 60–80% on mixed document collections.
Preserving Document Structure
Beyond extracting content, preserving document structure matters for downstream use cases like RAG and search. Position-aware extraction — tracking which page and approximately where on the page each piece of content appears — enables precise source citations in RAG applications. Docling provides bounding box information for extracted elements, which can be stored as metadata in the vector index and used to display the exact document location when retrieved. Heading hierarchy (H1, H2, H3) enables chunking strategies that keep related content together — splitting on section boundaries rather than arbitrary token counts produces significantly better RAG retrieval quality. Building this structural metadata into the extraction pipeline from the start is much easier than retrofitting it later.
Handling Password-Protected and Encrypted PDFs
Password-protected PDFs require decryption before processing. PyMuPDF handles this with the authenticate method when you have the password. For PDFs with permissions restrictions (print-allowed but copy-not-allowed), PyMuPDF’s page rendering still works — rendering to image does not trigger the copy restriction the way text extraction might, though the legality of this depends on your jurisdiction and use case. Heavily encrypted PDFs (where even rendering is blocked) require the password to decrypt fully. For enterprise document processing where password management is required, store passwords in a secrets manager keyed by document identifier and retrieve them programmatically before processing rather than embedding them in code.
Vision-based PDF processing is the right solution for document collections where text extraction fails or falls short. The pipeline is more expensive and slower than text extraction, but it handles the hard cases — scanned documents, complex layouts, figure-heavy content — that no text extraction library can manage. Start with a classification step to identify which PDFs actually need vision processing, use Docling or Marker for the bulk of them, and reserve direct vision LLM API calls for documents where those libraries need supplementation. The investment in a robust PDF processing pipeline pays for itself in downstream RAG quality and downstream data pipeline reliability.
Batch Processing Architecture
For enterprise-scale PDF processing, a robust batch architecture prevents partial failures from corrupting an entire run. The recommended pattern: a job queue (Celery, Redis Queue, or a cloud queue like SQS) accepts PDF processing tasks, worker processes pull tasks and process individual documents, results are written to a database keyed by document identifier with processing status (pending, processing, complete, failed), and a retry mechanism re-queues failed documents up to a configurable limit. This architecture handles API rate limits (by throttling queue consumption to stay within API limits), transient failures (by retrying with exponential backoff), and long documents (by tracking per-page progress so a failure mid-document does not require reprocessing from the start). For a collection of tens of thousands of documents, this architecture makes processing resumable and observable rather than an opaque batch job that fails silently.
Quality Assurance on Extracted Content
Vision-based PDF processing is not perfect — models occasionally misread text, skip content, or hallucinate elements, particularly on poor-quality scans. A lightweight QA layer catches systematic problems before extracted content enters production systems. Check completeness: compare the word count of extracted text to the expected word count based on page count and document type (a 10-page financial report should yield at least 1,500 words of text; extracting 200 words signals a problem). Check consistency: for documents with known structure (financial statements, standard forms), verify that expected fields are present. Check plausibility: run extracted text through a language detection model to verify it matches the expected language — extracting garbled characters is a common failure mode on non-Latin script documents. Flagging documents that fail these checks for manual review maintains extraction quality at scale without reviewing every document individually.