How to Analyse Charts with a Vision LLM

Analysing charts with a vision LLM means asking a multimodal model to read, interpret, and reason about visual data — bar charts, line graphs, scatter plots, pie charts, heatmaps — without converting them to numbers first. The use case is practical and common: extracting values from a chart image in a PDF report, comparing trends across multiple charts, identifying outliers in a scatter plot, or generating a written summary of what a chart shows. Vision LLMs handle this task reasonably well, but accuracy varies significantly by chart type, chart quality, and which model you use. Understanding the accuracy profile and the prompt patterns that improve it makes the difference between a reliable chart analysis tool and an unreliable one.

What Vision LLMs Can and Cannot Read

Vision LLMs are not OCR tools or data digitisers — they are pattern recognisers that extract meaning from visual representations. On clear, well-labelled charts with legible axis labels and values, current models (GPT-4o, Claude 3.5 Sonnet, InternVL2-26B) can read approximate values, identify trends, compare categories, and describe the chart accurately. On poor-quality charts — blurry images, unlabelled axes, overlapping elements, very small text — accuracy degrades significantly. The honest characterisation of what works: trend identification (is this going up, down, or flat) is reliable; relative comparison (which bar is tallest, which line peaked earlier) is reliable; exact numerical extraction (what is the precise value of this bar) is approximate, with errors of 5–15% common on bar charts and larger errors on pie charts. For use cases requiring precise values, supplementing vision LLM analysis with actual data extraction (if the chart data is accessible) is more reliable than relying solely on visual reading.

Chart Types and Their Accuracy Profiles

Not all charts are equally readable by vision LLMs. Bar charts are the most reliably analysed: clear height differences between bars are easy to compare, and labelled values (numbers printed on or above bars) are readable by OCR-capable vision models. Line charts are readable for trend direction and relative positioning, but exact values at specific points are less accurate. Scatter plots are well-handled for cluster identification and outlier detection, but specific coordinate values are unreliable. Pie charts are the least accurate for vision LLMs: accurately reading percentage values from arc lengths is difficult even for humans without labels, and models often produce approximations that sum to less or more than 100%. Heatmaps depend heavily on the colour scale — a well-labelled continuous colour scale enables reasonable value estimation; unlabelled or perceptually ambiguous colour scales (rainbow scales) produce poor results. Stacked bar and area charts are often misread — models struggle to correctly attribute values to individual stacked segments.

Prompting Strategies for Chart Analysis

The quality of chart analysis depends heavily on how you prompt the model. Generic prompts (“describe this chart”) produce generic descriptions. Specific prompts targeting the information you actually need produce much more useful output. For value extraction: “Read the y-axis values for each bar in this chart. List them as: [category name]: [value]. If a value is not clearly labelled, estimate it from the bar height relative to the axis scale.” For trend analysis: “Describe the trend shown in this line chart over time. Identify the overall direction, any notable peaks or troughs, and approximately when they occurred.” For comparison: “Compare the values across the groups shown in this chart. Which group has the highest value? What is the approximate ratio between the largest and smallest values?” For summary generation: “Provide a three-sentence summary of what this chart shows, including the main finding, any notable patterns, and any important caveats visible in the chart.” Explicit, targeted prompts consistently outperform vague ones because they direct the model’s attention to the specific visual elements relevant to the analysis task.

Figure 1 — Chart type accuracy guide for vision LLMs

Chart type Trend reading Value accuracy Notes Bar chartExcellentGood (±5–10%)Best with labelled values Line chartExcellentModerate (±10%)Trends reliable; points less so Scatter plotGoodPoor (coordinates)Clusters/outliers reliable Pie chartModeratePoor (±15–25%)Use only with labels HeatmapGoodModerateDepends on colour scale quality Accuracy with GPT-4o / InternVL2-26B on clean, well-labelled charts

Model Choice for Chart Analysis

Not all vision models are equally capable at chart analysis, and the differences are larger than on natural image tasks. GPT-4o and Claude 3.5/3.7 Sonnet are the most accurate for chart reading among commercial models — both have been fine-tuned or trained on datasets that include data visualisations, and both show strong OCR capability for reading labelled values. Among open-source models, InternVL2-26B and InternVL2-76B are the strongest for document and chart analysis; InternVL2-8B handles clear charts well but struggles with small text and complex multi-series charts. LLaVA models are weaker on chart analysis than InternVL2 at comparable parameter counts — the OCRBench gap (794 vs 532) translates directly to chart reading accuracy. For local deployment where chart analysis quality matters, InternVL2-26B (requiring approximately 52GB VRAM at 4-bit) is the recommended minimum; for constrained hardware, InternVL2-8B handles most clearly-labelled charts acceptably.

Image Quality Preprocessing

Chart analysis quality is significantly affected by input image quality. Several preprocessing steps improve results. Resolution: charts below 800×600 pixels are often too low-resolution for accurate text reading; upscale with a bicubic or super-resolution filter before passing to the model. Contrast enhancement: low-contrast charts (light-coloured bars on white backgrounds, muted colour schemes) benefit from contrast adjustment. Cropping: pass only the chart region rather than the full page — removing surrounding text and headers reduces the amount of visual information the model processes and focuses attention on the chart. For PDFs, rendering at 150–200 DPI before extracting charts provides sufficient resolution for most analysis tasks. For scanned documents, apply deskewing and denoising before chart analysis — crooked or noisy scans significantly degrade OCR-dependent chart reading.

When to Use Vision LLMs vs Data Extraction

Vision LLMs are not the right tool for every chart analysis task. If the underlying data is accessible — the chart was generated from a spreadsheet or database that you can query — extract the data directly rather than reading values from the chart image. If you need precise values for downstream calculations, the approximation error from visual reading (5–15% for bar charts) compounds and produces unreliable results. Vision LLM chart analysis is the right tool when: the data is not otherwise accessible (charts embedded in PDFs from external sources); you need a natural language summary rather than precise values; you need to answer qualitative questions about trends or comparisons; or you need to process large volumes of charts quickly where human review is impractical. For enterprise document intelligence applications where charts appear in analyst reports, financial filings, or technical specifications, vision LLM chart analysis — even with its accuracy limitations — extracts value that text-only approaches miss entirely.

Structured Output for Chart Data

For chart analysis at scale, prompting the model to return structured output (JSON) rather than prose makes downstream processing reliable. Define the expected output schema: for a bar chart, a list of objects with category and value fields; for a line chart, a list of time-point and value pairs; for a comparison chart, a structured object with the key insight and supporting evidence. Combining a JSON schema in the prompt with structured output enforcement (function calling or structured output APIs in OpenAI/Anthropic) reduces parsing errors and makes chart analysis results directly usable in data pipelines. Validation against the schema catches cases where the model produced an estimate it was not confident in, which can be flagged for human review rather than silently passed downstream with a potentially large error.

Chart analysis with vision LLMs delivers genuine value for document intelligence applications — it makes visual data accessible to automated pipelines that previously could only process text. The accuracy is sufficient for trend detection, qualitative comparison, and summary generation on clean, well-labelled charts. For precise value extraction, supplement with validation steps or data extraction where possible. Use InternVL2 or GPT-4o-class models, preprocess images to adequate resolution, and prompt specifically for the information you need rather than asking for a generic description. The combination of a capable vision model and a targeted prompt is what separates reliable chart analysis from inconsistent results.

One practical workflow pattern worth noting: run a fast, cheap model (Haiku, Gemini Flash) as a first pass to classify which charts are high-value and which are decorative or redundant, then route only the high-value ones to a more capable model for detailed analysis. This tiered approach reduces cost significantly on document collections with mixed chart quality without sacrificing accuracy on the charts that matter most for the downstream task.

Leave a Comment