Can an AI model read a PDF?
The registry verdict table for reading a PDF
Of the ten registry configurations assessed against this family's own canonical task, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.
| Model | Manufacturer | Training cutoff | Verdict | Score | Harmony |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google DeepMind | 2025-01 | Inauspicious | 1 of 7 | -4 |
| DeepSeek V4 Pro | DeepSeek | 2025-12 | Favourable | 5 of 7 | -1 |
| Claude Fable 5 | Anthropic | 2026-01 | Highly auspicious | 7 of 7 | +5 |
| Claude Sonnet 5 | Anthropic | 2026-01 | Neutral | 4 of 7 | -3 |
| GPT-5.6 Sol | OpenAI | 2026-02 | Unfavourable | 2 of 7 | +1 |
| GPT-5.6 Terra | OpenAI | 2026-02 | Unfavourable | 2 of 7 | -1 |
| Grok 4.6 | xAI | 2026-02 | Auspicious | 6 of 7 | -1 |
| Kimi K3 | Moonshot AI | 2026-03 | Inauspicious | 1 of 7 | -1 |
| Qwen3.8-Max | Alibaba Cloud | 2026-03 | Guarded | 3 of 7 | 0 |
| Claude Opus 5 | Anthropic | 2026-05 | Auspicious | 6 of 7 | +2 |
The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.
About this capability question
Reading a PDF is, by the Registry's own measurement, the family with the least purchased search-volume evidence behind it and simultaneously one of the most consistently recurring in raw harvested search behaviour — eight distinct phrasings, from 'read pdfs' to the oddly emphatic 'not read pdfs,' recurring behind six of the eight brand prefixes tracked. Purchased keyword data measures what has already been priced for advertising; it does not measure what people actually type, and this family is the clearest evidence of the gap between the two.
The canonical brief assessed here is specific rather than generic: a signed contract, with payment terms and a renewal date extracted from it. Specificity is deliberate — a registry-canonical task string has to name something concrete enough to hash consistently, and 'read a PDF' alone does not.
No actual contract sits behind that brief. It names a document shape — signed, with payment terms and a renewal date — without naming any real agreement, real party, or real figure, because the wording is what gets hashed, not any file.
How the ten configurations read reading a PDF
Information retrieval is the most heavily represented task class in this wave, with five families sharing it, so the dominant-house and dominant-element reading below is identical across all five. What separates this page from its retrieval siblings is its own canonical brief and the verdict distribution that brief produces.
Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to reading a PDF: reading information retrieval through a dominant house or a dominant element, 2 of the ten configurations; and, from this family's own canonical task, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten in the distribution below.
| Published value | Verdict | Configurations |
|---|---|---|
| 1 of 7 | Inauspicious | 2 of the ten |
| 2 of 7 | Unfavourable | 2 of the ten |
| 3 of 7 | Guarded | 1 of the ten |
| 4 of 7 | Neutral | 1 of the ten |
| 5 of 7 | Favourable | 1 of the ten |
| 6 of 7 | Auspicious | 2 of the ten |
| 7 of 7 | Highly auspicious | 1 of the ten |
Why reading a pdf is classified as Information retrieval
Extracting named terms from an existing document is an act of retrieval — locating and reporting on content that already exists in a fixed form — rather than a transformation or a new artefact.
On the order form's task class field, select Information retrieval. That class's own risk-factor page is /tasks/retrieval.
Extracting terms from a document carries the risk of missing a clause or misreading a figure inside it — errors invisible in the extracted output unless someone checks it against the original contract. 4 risk factors are tracked for information retrieval tasks generally; an assessment for reading a PDF reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.
DeepSeek V4 Pro, derivation shown working
DeepSeek V4 Pro's permanent chart never reads a task. Its chart_seed is the
SHA-256 digest of the string "DeepSeek V4 Pro|2025-12|0.700"
— d37032cc2be5b9e0…, the first sixteen of sixty-four hex
characters — and that seed alone fixes an ascendant of Scorpio, a
ruling planet of Sun, and a harmony of -1.
The same chart appears, unchanged, on every task-class page this configuration is
assessed against.
What reads reading a PDF specifically is this family's own canonical
string — never stored, only hashed, to f4c915f211f6a01f… —
combined with the chart_seed and the task class (retrieval) to form
task_seed. Its first byte, taken modulo seven and offset by that harmony of
-1, is the verdict's internal base, clamped to the
0–6 range and published one higher on the 1–7 scale /method documents in full.
This configuration's harmony is not enough, for this brief, to push the internal score against either bound: the published result, 5 of 7 — Favourable, lands 4 steps clear of the floor and 2 clear of the ceiling.
What this page does not claim about reading a PDF
Whether a given contract is three pages or three hundred, typed or scanned, is not information the canonical brief encodes, and none of it enters the computation below. The brief names the extraction task, not a document's format. The Registry does not test models and does not measure a success rate for reading a PDF or any other task; what it does is compute, and the table at the top of this page is what it computed.
A buyer's own reading a PDF task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.
Related surfaces
← All capability questions · The AFR-1 method · Information retrieval task class
Related capability questions: Read an Excel file · Open a link · Open a ZIP file
Questions about reading a PDF
- Can an AI model read a PDF?
- Across the ten registry configurations at their own registry-default settings, no single word dominates: Inauspicious, Unfavourable, Auspicious each account for 2 of the ten for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Inauspicious and the newest returns Auspicious; the full ten-row table above lists the exact verdict for every configuration in between.
- Why is reading a PDF classified as Information retrieval rather than a different task class?
- Extracting named terms from an existing document is an act of retrieval — locating and reporting on content that already exists in a fixed form — rather than a transformation or a new artefact.
- Why does this page cite no external search-volume figure the way some other Registry pages do?
- Because none exists in the purchased keyword data for this specific family — a gap the Registry states rather than papers over with an estimate. The underlying request is nonetheless well evidenced in raw harvested search phrasing, described in the paragraph above.