Can an AI model transcribe audio?
The registry verdict table for transcribing audio
Of the ten registry configurations assessed against this family's own canonical task, no single word dominates: Inauspicious, Guarded each account for 3 of the ten, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.
| Model | Manufacturer | Training cutoff | Verdict | Score | Harmony |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google DeepMind | 2025-01 | Inauspicious | 1 of 7 | -4 |
| DeepSeek V4 Pro | DeepSeek | 2025-12 | Guarded | 3 of 7 | -1 |
| Claude Fable 5 | Anthropic | 2026-01 | Auspicious | 6 of 7 | +5 |
| Claude Sonnet 5 | Anthropic | 2026-01 | Inauspicious | 1 of 7 | -3 |
| GPT-5.6 Sol | OpenAI | 2026-02 | Guarded | 3 of 7 | +1 |
| GPT-5.6 Terra | OpenAI | 2026-02 | Unfavourable | 2 of 7 | -1 |
| Grok 4.6 | xAI | 2026-02 | Inauspicious | 1 of 7 | -1 |
| Kimi K3 | Moonshot AI | 2026-03 | Auspicious | 6 of 7 | -1 |
| Qwen3.8-Max | Alibaba Cloud | 2026-03 | Guarded | 3 of 7 | 0 |
| Claude Opus 5 | Anthropic | 2026-05 | Highly auspicious | 7 of 7 | +2 |
The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.
About this capability question
Transcribing audio runs in the direction opposite every media family in this wave: sound goes in, and text — a written record of an existing recording, not a new artefact — comes out. The canonical version assessed here is a recorded customer support call, timestamped on the way to text, which is a common enough operational shape that the request recurs, largely unchanged, behind all eight brand prefixes tracked.
Because the output is a modality change rather than a new artefact, the Registry's own classification places this family with data processing rather than with media generation — a distinction worth stating plainly before the verdict table, since a superficial reading of 'audio in, something out' could suggest otherwise.
The canonical brief published below names a customer support call specifically, rather than audio in general, for the same reason every family on this registry commits to a concrete wording: a task string that could mean several different things would hash inconsistently against itself in every reader's head, even though the underlying hash is perfectly consistent in the engine.
How the ten configurations read transcribing audio
Data processing is one of the classes the Registry's correspondence tables do assign a dominant house and element to, so the ten configurations below carry a genuine class-affinity reading rather than the class-neutral result media generation gets. That reading, and the verdict distribution beside it, are both specific to this family's own computed chain.
Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to transcribing audio: reading data processing through a dominant house or a dominant element, 3 of the ten configurations; and, from this family's own canonical task, no single word dominates: Inauspicious, Guarded each account for 3 of the ten in the distribution below.
| Published value | Verdict | Configurations |
|---|---|---|
| 1 of 7 | Inauspicious | 3 of the ten |
| 2 of 7 | Unfavourable | 1 of the ten |
| 3 of 7 | Guarded | 3 of the ten |
| 4 of 7 | Neutral | 0 of the ten |
| 5 of 7 | Favourable | 0 of the ten |
| 6 of 7 | Auspicious | 2 of the ten |
| 7 of 7 | Highly auspicious | 1 of the ten |
Why transcribing audio is classified as Data processing
Transcription converts an existing recording into an existing format — text — rather than manufacturing something that did not exist before the brief. That is a transformation, and the Registry's data processing class is defined around exactly that operation.
On the order form's task class field, select Data processing. That class's own risk-factor page is /tasks/analysis.
Transcription risk is about fidelity to a recording that already exists — a wrong word, a missed timestamp, a garbled name — the opposite failure shape from a media-generation brief with no source to check against. 4 risk factors are tracked for data processing tasks generally; an assessment for transcribing audio reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.
Claude Sonnet 5, derivation shown working
Claude Sonnet 5's permanent chart never reads a task. Its chart_seed is the
SHA-256 digest of the string "Claude Sonnet 5|2026-01|0.700"
— fc36941ed9d3b87c…, the first sixteen of sixty-four hex
characters — and that seed alone fixes an ascendant of Aries, a
ruling planet of Venus, and a harmony of -3.
The same chart appears, unchanged, on every task-class page this configuration is
assessed against.
What reads transcribing audio specifically is this family's own canonical
string — never stored, only hashed, to 0d6721764069c751… —
combined with the chart_seed and the task class (analysis) to form
task_seed. Its first byte, taken modulo seven and offset by that harmony of
-3, is the verdict's internal base, clamped to the
0–6 range and published one higher on the 1–7 scale /method documents in full.
Nothing left to clamp toward here: for this brief, this configuration's internal score sits at the floor, published as 1 of 7 — Inauspicious.
What this page does not claim about transcribing audio
A timestamp accurate to the second and a timestamp accurate to the minute are both, for the purposes of this table, simply 'transcribe a call' — the canonical brief fixes wording, not precision, and the method has no finer-grained input to read. The Registry does not test models and does not measure a success rate for transcribing audio or any other task; what it does is compute, and the table at the top of this page is what it computed.
A buyer's own transcribing audio task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.
Related surfaces
← All capability questions · The AFR-1 method · Data processing task class
Related capability questions: Summarize a YouTube video · Read a PDF · Make a video
Questions about transcribing audio
- Can an AI model transcribe audio?
- Across the ten registry configurations at their own registry-default settings, no single word dominates: Inauspicious, Guarded each account for 3 of the ten for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Inauspicious and the newest returns Highly auspicious; the full ten-row table above lists the exact verdict for every configuration in between.
- Why is transcribing audio classified as Data processing rather than a different task class?
- Transcription converts an existing recording into an existing format — text — rather than manufacturing something that did not exist before the brief. That is a transformation, and the Registry's data processing class is defined around exactly that operation.
- Does a lower verdict for a configuration mean transcription accuracy would be worse on that configuration?
- No. The verdict is the output of a deterministic method applied to a fixed configuration and a fixed canonical brief; it carries no information about word error rate, timestamp precision, or any other property of an actual transcript, because none of those were measured to produce it.