How to Extract Text from Images with an LLM

Extracting text from images with a language model means using a vision-capable LLM as an OCR engine — passing an image to the model and asking it to return the text it contains. For many use cases this works better than traditional OCR: it handles handwriting, low-quality scans, unusual fonts, mixed languages, and complex layouts more robustly than rule-based OCR systems, and it can simultaneously understand the text it extracts, not just transcribe it. Understanding when to use a vision LLM for text extraction, which models to use, and how to prompt them effectively makes the difference between a reliable extraction pipeline and an inconsistent one.

When Vision LLMs Beat Traditional OCR

Traditional OCR engines (Tesseract, PaddleOCR, AWS Textract) are optimised for clean, typed text on good-quality scans. They are fast, cheap, and highly accurate in that regime. Vision LLMs outperform them in several specific scenarios: handwritten text, where the character recognition problem is significantly harder and traditional OCR degrades sharply; mixed-language documents, where OCR engines need separate language models and struggle with code-switching; low-quality images — blurry, low-contrast, skewed, or heavily compressed — where vision LLMs’ contextual understanding helps fill in ambiguous characters; complex layouts with tables, forms, or mixed text and figures where traditional OCR extracts text in wrong reading order; and mathematical notation or domain-specific symbols that are outside standard OCR character sets. For plain typed text on a clean document, traditional OCR remains faster and cheaper. For any of the above cases, a vision LLM is the right tool.

Model Selection for Text Extraction

OCR capability varies significantly across vision models, and it is worth knowing the ranking before selecting a model. InternVL2 family models consistently lead open-source benchmarks on OCRBench — InternVL2-8B scores approximately 794 on OCRBench versus LLaVA-1.6’s 532. Among commercial models, GPT-4o and Claude 3.5 Sonnet are both strong for document text extraction, with GPT-4o showing a slight edge on low-quality scans. For local deployment, InternVL2-8B handles most business document scenarios reliably; InternVL2-26B (requiring approximately 52GB VRAM at 4-bit) adds capability on handwritten text and very complex layouts. Moondream2, a 2B model designed for efficiency, handles simple typed text extraction well and runs on 4GB VRAM — useful for constrained deployments where document quality is consistently good. For mixed handwritten and typed documents, a model with strong OCR training (InternVL2, GPT-4o) is necessary; Moondream and LLaVA models are insufficient for challenging handwriting.

Prompting for Clean Extraction

The prompt significantly affects extraction quality. A generic “what does this image say?” produces less reliable output than a task-specific extraction prompt. For complete document transcription: “Transcribe all text visible in this image exactly as it appears. Preserve line breaks and spacing. Do not add any commentary or interpretation — only the text as written.” For structured forms: “Extract the field labels and their values from this form image. Return them as a JSON object where keys are field labels and values are the filled-in content.” For receipts or invoices: “Extract the following fields from this receipt: vendor name, date, line items (description and price), subtotal, tax, and total. Return as structured JSON.” For mixed content: “Transcribe only the text in this image. Ignore any graphics, logos, or decorative elements. If text is crossed out or struck through, include it with a note that it is struck through.” Specific, directive prompts reduce hallucination risk and improve structural accuracy.

Figure 1 — Text extraction approaches: when to use each

Document type Best approach Notes Clean typed text, PDFTraditional OCRFaster and cheaper; use pdfplumber/pypdf Handwritten notesVision LLMInternVL2-8B+ or GPT-4o Scanned forms / receiptsVision LLMStructured output prompt; GPT-4o reliable Complex mixed layoutVision LLMPreserves reading order; InternVL2-26B best Math / symbolsVision LLMLaTeX output prompt for equations

Handling Multi-Page Documents

For documents longer than a single image, process each page separately and concatenate results. The recommended pipeline: convert the PDF to images (PyMuPDF renders pages as PNG at configurable DPI — 150 DPI is sufficient for most text extraction), process each page image through the vision LLM, and assemble the extracted text with page markers. Page-by-page processing keeps each API call within model input limits and allows parallel processing of pages for throughput. For very long documents where API cost is a concern, add a triage step: run a fast lightweight model (Claude Haiku, Gemini Flash) to identify pages likely to contain valuable text versus cover pages, blank pages, or purely decorative content, then route only the high-value pages to the more expensive accurate model.

Table Extraction

Tables in images are a particularly valuable extraction target — they contain structured data that is genuinely hard to extract with traditional OCR (which reads cell by cell without understanding table structure). Vision LLMs handle tables well when prompted specifically: “Extract the table from this image. Return the content as a markdown table, preserving all rows and columns. Include column headers if visible.” For tables that need to be used programmatically, request JSON output: “Extract the table as a JSON array of objects, where each object represents a row with column header names as keys.” GPT-4o and InternVL2-26B are the most reliable for complex tables with merged cells or nested headers; simpler tables are handled well by smaller models. For financial tables with many decimal values, verify a sample of extracted figures against the source image — numeric OCR errors (1 vs 7, 0 vs 6) are the most consequential failure mode.

Mathematical and Scientific Notation

Extracting mathematical expressions from images requires a model that can produce LaTeX output. For equations: “Transcribe the mathematical expression in this image as LaTeX. Return only the LaTeX code, nothing else.” GPT-4o handles standard mathematical notation reliably. For handwritten equations, which are common in notes, textbooks, and exam answers, accuracy depends heavily on the clarity of handwriting — printed handwriting is handled well, cursive or informal notation is less reliable. Pix2Tex (LaTeX-OCR) is a specialised open-source model trained specifically for mathematical expression recognition and outperforms general vision LLMs on equations extracted from printed text; for handwritten equations, general vision LLMs remain more robust. A hybrid approach — use Pix2Tex for detected equation regions and a general vision LLM for surrounding text — maximises accuracy on mixed mathematical documents.

Confidence and Validation

Vision LLMs occasionally hallucinate text that is not in the image, particularly when image quality is low or text is partially obscured. For high-stakes extraction (legal documents, financial figures, medical records), add a validation step: ask the model to rate its confidence in the extraction, flag low-confidence regions, or re-extract a second time and compare the two extractions for consistency. When two independent extractions disagree on a specific field, flag that field for human review rather than selecting one output arbitrarily. For numerical extraction specifically, cross-validate key figures: if a receipt shows a subtotal, tax, and total, verify that subtotal + tax ≈ total as a consistency check. Automated consistency checks catch a significant fraction of extraction errors without requiring manual review of every document.

Vision LLM text extraction is the right tool for challenging document types — handwriting, complex layouts, low-quality scans, mathematical notation — where traditional OCR fails. For clean typed text, traditional OCR remains faster and cheaper. The practical approach is a routing layer: classify document type and quality, send clear typed text to a traditional OCR pipeline and everything else to a vision LLM. This hybrid approach optimises both cost and accuracy across a mixed document collection without sacrificing performance on either document type.

Image Preprocessing for Better Extraction

Input image quality directly affects extraction accuracy. Several preprocessing steps reliably improve results. Deskewing: images captured by phone cameras are often slightly tilted — automatic deskewing (using OpenCV’s warpAffine or the deskew Python library) corrects this before passing to the model. Binarisation: for low-contrast scanned documents, converting to black-and-white with adaptive thresholding increases text visibility, particularly for faded ink. Resolution normalisation: images below 100 DPI produce degraded OCR results with any model — upscale to at least 150 DPI using bicubic interpolation. Noise removal: salt-and-pepper noise from scanner artifacts degrades recognition; a median filter removes it without blurring text edges. For mobile-captured documents specifically — photos of whiteboards, receipts, or printed pages — these preprocessing steps can reduce error rates by 20–40% compared to passing raw phone photos directly to the model. The OpenCV library handles all of these steps efficiently in a preprocessing pipeline that adds negligible latency to the overall extraction workflow.

Leave a Comment