Can an AI model read a PDF?

The registry verdict table for reading a PDF

Of the ten registry configurations assessed against this family's own canonical task, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.

Registry-computed verdict, score, and composite weight for reading a PDF, one row per registry model configuration
Model Manufacturer Training cutoff Verdict Score Harmony
Gemini 3.7 Flash Google DeepMind 2025-01 Inauspicious 1 of 7 -4
DeepSeek V4 Pro DeepSeek 2025-12 Favourable 5 of 7 -1
Claude Fable 5 Anthropic 2026-01 Highly auspicious 7 of 7 +5
Claude Sonnet 5 Anthropic 2026-01 Neutral 4 of 7 -3
GPT-5.6 Sol OpenAI 2026-02 Unfavourable 2 of 7 +1
GPT-5.6 Terra OpenAI 2026-02 Unfavourable 2 of 7 -1
Grok 4.6 xAI 2026-02 Auspicious 6 of 7 -1
Kimi K3 Moonshot AI 2026-03 Inauspicious 1 of 7 -1
Qwen3.8-Max Alibaba Cloud 2026-03 Guarded 3 of 7 0
Claude Opus 5 Anthropic 2026-05 Auspicious 6 of 7 +2

The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.

About this capability question

Reading a PDF is, by the Registry's own measurement, the family with the least purchased search-volume evidence behind it and simultaneously one of the most consistently recurring in raw harvested search behaviour — eight distinct phrasings, from 'read pdfs' to the oddly emphatic 'not read pdfs,' recurring behind six of the eight brand prefixes tracked. Purchased keyword data measures what has already been priced for advertising; it does not measure what people actually type, and this family is the clearest evidence of the gap between the two.

The canonical brief assessed here is specific rather than generic: a signed contract, with payment terms and a renewal date extracted from it. Specificity is deliberate — a registry-canonical task string has to name something concrete enough to hash consistently, and 'read a PDF' alone does not.

No actual contract sits behind that brief. It names a document shape — signed, with payment terms and a renewal date — without naming any real agreement, real party, or real figure, because the wording is what gets hashed, not any file.

How the ten configurations read reading a PDF

Information retrieval is the most heavily represented task class in this wave, with five families sharing it, so the dominant-house and dominant-element reading below is identical across all five. What separates this page from its retrieval siblings is its own canonical brief and the verdict distribution that brief produces.

Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to reading a PDF: reading information retrieval through a dominant house or a dominant element, 2 of the ten configurations; and, from this family's own canonical task, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten in the distribution below.

Verdict distribution for reading a PDF, ten registry configurations, fixed 1–7 scale order
Published valueVerdictConfigurations
1 of 7Inauspicious2 of the ten
2 of 7Unfavourable2 of the ten
3 of 7Guarded1 of the ten
4 of 7Neutral1 of the ten
5 of 7Favourable1 of the ten
6 of 7Auspicious2 of the ten
7 of 7Highly auspicious1 of the ten

Why reading a pdf is classified as Information retrieval

Extracting named terms from an existing document is an act of retrieval — locating and reporting on content that already exists in a fixed form — rather than a transformation or a new artefact.

On the order form's task class field, select Information retrieval. That class's own risk-factor page is /tasks/retrieval.

Extracting terms from a document carries the risk of missing a clause or misreading a figure inside it — errors invisible in the extracted output unless someone checks it against the original contract. 4 risk factors are tracked for information retrieval tasks generally; an assessment for reading a PDF reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.

DeepSeek V4 Pro, derivation shown working

DeepSeek V4 Pro's permanent chart never reads a task. Its chart_seed is the SHA-256 digest of the string "DeepSeek V4 Pro|2025-12|0.700" — d37032cc2be5b9e0…, the first sixteen of sixty-four hex characters — and that seed alone fixes an ascendant of Scorpio, a ruling planet of Sun, and a harmony of -1. The same chart appears, unchanged, on every task-class page this configuration is assessed against.

What reads reading a PDF specifically is this family's own canonical string — never stored, only hashed, to f4c915f211f6a01f… — combined with the chart_seed and the task class (retrieval) to form task_seed. Its first byte, taken modulo seven and offset by that harmony of -1, is the verdict's internal base, clamped to the 0–6 range and published one higher on the 1–7 scale /method documents in full. This configuration's harmony is not enough, for this brief, to push the internal score against either bound: the published result, 5 of 7 — Favourable, lands 4 steps clear of the floor and 2 clear of the ceiling.

What this page does not claim about reading a PDF

Whether a given contract is three pages or three hundred, typed or scanned, is not information the canonical brief encodes, and none of it enters the computation below. The brief names the extraction task, not a document's format. The Registry does not test models and does not measure a success rate for reading a PDF or any other task; what it does is compute, and the table at the top of this page is what it computed.

A buyer's own reading a PDF task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.

Get Assessment

Related surfaces

← All capability questions · The AFR-1 method · Information retrieval task class

Related capability questions: Read an Excel file · Open a link · Open a ZIP file

Questions about reading a PDF

Can an AI model read a PDF?
Across the ten registry configurations at their own registry-default settings, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Inauspicious and the newest returns Auspicious; the full ten-row table above lists the exact verdict for every configuration in between.
Why is reading a PDF classified as Information retrieval rather than a different task class?
Extracting named terms from an existing document is an act of retrieval — locating and reporting on content that already exists in a fixed form — rather than a transformation or a new artefact.
Why does this page cite no external search-volume figure the way some other Registry pages do?
Because none exists in the purchased keyword data for this specific family — a gap the Registry states rather than papers over with an estimate. The underlying request is nonetheless well evidenced in raw harvested search phrasing, described in the paragraph above.