Can an AI model write code?
The registry verdict table for writing code
Of the ten registry configurations assessed against this family's own canonical task, Inauspicious for 3 of them — the modal result, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.
| Model | Manufacturer | Training cutoff | Verdict | Score | Harmony |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google DeepMind | 2025-01 | Inauspicious | 1 of 7 | -4 |
| DeepSeek V4 Pro | DeepSeek | 2025-12 | Inauspicious | 1 of 7 | -1 |
| Claude Fable 5 | Anthropic | 2026-01 | Highly auspicious | 7 of 7 | +5 |
| Claude Sonnet 5 | Anthropic | 2026-01 | Inauspicious | 1 of 7 | -3 |
| GPT-5.6 Sol | OpenAI | 2026-02 | Auspicious | 6 of 7 | +1 |
| GPT-5.6 Terra | OpenAI | 2026-02 | Guarded | 3 of 7 | -1 |
| Grok 4.6 | xAI | 2026-02 | Unfavourable | 2 of 7 | -1 |
| Kimi K3 | Moonshot AI | 2026-03 | Neutral | 4 of 7 | -1 |
| Qwen3.8-Max | Alibaba Cloud | 2026-03 | Unfavourable | 2 of 7 | 0 |
| Claude Opus 5 | Anthropic | 2026-05 | Highly auspicious | 7 of 7 | +2 |
The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.
About this capability question
Writing code recurred behind seven of the eight brand prefixes this registry tracked — every brand except the one built specifically for coding, an absence worth noting rather than smoothing over: a purpose-built coding assistant is evidently not who most people are asking this question of. The canonical brief here is a small, self-contained unit: a Python function that validates and parses a CSV file, delivered with its own tests.
Unit tests are named explicitly in that brief, and they matter to how this class is framed generally: a test suite passing is evidence about the paths it exercises, not a guarantee about the paths it does not, a distinction the classification below is built around.
The verdict table below answers for this exact brief and no other. A buyer submitting a real function signature, a real CSV shape, or a real test framework is answered by a separate order, computed against what they actually submit rather than against the generic CSV-parsing brief shown here.
Kimi K3, derivation shown working
Kimi K3's permanent chart never reads a task. Its chart_seed is the
SHA-256 digest of the string "Kimi K3|2026-03|0.700"
— 66b027590ef48c98…, the first sixteen of sixty-four hex
characters — and that seed alone fixes an ascendant of Libra, a
ruling planet of Moon, and a harmony of -1.
The same chart appears, unchanged, on every task-class page this configuration is
assessed against.
What reads writing code specifically is this family's own canonical
string — never stored, only hashed, to e58b203026a3e8cf… —
combined with the chart_seed and the task class (code) to form
task_seed. Its first byte, taken modulo seven and offset by that harmony of
-1, is the verdict's internal base, clamped to the
0–6 range and published one higher on the 1–7 scale /method documents in full.
Neither bound of the 0–6 range is reached for this brief: 4 of 7 — Neutral — sits 3 above the floor and 3 below the ceiling, an unclamped result.
How the ten configurations read writing code
This is the wave's other same-class collision: build-a-website shares code modification with this family, so the dominant-house and dominant-element reading below repeats between the two pages exactly. The verdict distribution is where the two pages part company.
Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to writing code: reading code modification through a dominant house or a dominant element, 5 of the ten configurations; and, from this family's own canonical task, Inauspicious for 3 of them — the modal result in the distribution below.
| Published value | Verdict | Configurations |
|---|---|---|
| 1 of 7 | Inauspicious | 3 of the ten |
| 2 of 7 | Unfavourable | 2 of the ten |
| 3 of 7 | Guarded | 1 of the ten |
| 4 of 7 | Neutral | 1 of the ten |
| 5 of 7 | Favourable | 0 of the ten |
| 6 of 7 | Auspicious | 1 of the ten |
| 7 of 7 | Highly auspicious | 2 of the ten |
Why writing code is classified as Code modification
The task both produces new source code and asserts, through its own tests, something about that code's correctness — squarely the code modification class, defined around exactly this combination of writing and verifying.
On the order form's task class field, select Code modification. That class's own risk-factor page is /tasks/code.
A written function's risk is that a passing test suite proves nothing about the paths it never exercised — a quiet, well-formatted failure rather than a loud one, which is the shape this class exists to track. 4 risk factors are tracked for code modification tasks generally; an assessment for writing code reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.
What this page does not claim about writing code
Which CSV dialect, which validation rules, and how many unit tests would be enough are all questions a real implementation answers and the canonical brief leaves open. The table below reads the brief as worded, not as any particular implementation would resolve it. The Registry does not test models and does not measure a success rate for writing code or any other task; what it does is compute, and the table at the top of this page is what it computed.
A buyer's own writing code task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.
Related surfaces
← All capability questions · The AFR-1 method · Code modification task class
Related capability questions: Build a website · Open a link · Send an email
Questions about writing code
- Can an AI model write code?
- Across the ten registry configurations at their own registry-default settings, Inauspicious for 3 of them — the modal result for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Inauspicious and the newest returns Highly auspicious; the full ten-row table above lists the exact verdict for every configuration in between.
- Why is writing code classified as Code modification rather than a different task class?
- The task both produces new source code and asserts, through its own tests, something about that code's correctness — squarely the code modification class, defined around exactly this combination of writing and verifying.
- Would a different programming language change which task class this family is assessed under?
- No. Code modification covers writing, editing, refactoring, and testing code in any language; the canonical brief specifies Python because a brief needs one concrete wording, not because the classification is language-specific.