Can an AI model write code?

The registry verdict table for writing code

Of the ten registry configurations assessed against this family's own canonical task, Inauspicious for 3 of them — the modal result, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.

Registry-computed verdict, score, and composite weight for writing code, one row per registry model configuration
Model Manufacturer Training cutoff Verdict Score Harmony
Gemini 3.7 Flash Google DeepMind 2025-01 Inauspicious 1 of 7 -4
DeepSeek V4 Pro DeepSeek 2025-12 Inauspicious 1 of 7 -1
Claude Fable 5 Anthropic 2026-01 Highly auspicious 7 of 7 +5
Claude Sonnet 5 Anthropic 2026-01 Inauspicious 1 of 7 -3
GPT-5.6 Sol OpenAI 2026-02 Auspicious 6 of 7 +1
GPT-5.6 Terra OpenAI 2026-02 Guarded 3 of 7 -1
Grok 4.6 xAI 2026-02 Unfavourable 2 of 7 -1
Kimi K3 Moonshot AI 2026-03 Neutral 4 of 7 -1
Qwen3.8-Max Alibaba Cloud 2026-03 Unfavourable 2 of 7 0
Claude Opus 5 Anthropic 2026-05 Highly auspicious 7 of 7 +2

The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.

About this capability question

Writing code recurred behind seven of the eight brand prefixes this registry tracked — every brand except the one built specifically for coding, an absence worth noting rather than smoothing over: a purpose-built coding assistant is evidently not who most people are asking this question of. The canonical brief here is a small, self-contained unit: a Python function that validates and parses a CSV file, delivered with its own tests.

Unit tests are named explicitly in that brief, and they matter to how this class is framed generally: a test suite passing is evidence about the paths it exercises, not a guarantee about the paths it does not, a distinction the classification below is built around.

The verdict table below answers for this exact brief and no other. A buyer submitting a real function signature, a real CSV shape, or a real test framework is answered by a separate order, computed against what they actually submit rather than against the generic CSV-parsing brief shown here.

Kimi K3, derivation shown working

Kimi K3's permanent chart never reads a task. Its chart_seed is the SHA-256 digest of the string "Kimi K3|2026-03|0.700" — 66b027590ef48c98…, the first sixteen of sixty-four hex characters — and that seed alone fixes an ascendant of Libra, a ruling planet of Moon, and a harmony of -1. The same chart appears, unchanged, on every task-class page this configuration is assessed against.

What reads writing code specifically is this family's own canonical string — never stored, only hashed, to e58b203026a3e8cf… — combined with the chart_seed and the task class (code) to form task_seed. Its first byte, taken modulo seven and offset by that harmony of -1, is the verdict's internal base, clamped to the 0–6 range and published one higher on the 1–7 scale /method documents in full. Neither bound of the 0–6 range is reached for this brief: 4 of 7 — Neutral — sits 3 above the floor and 3 below the ceiling, an unclamped result.

How the ten configurations read writing code

This is the wave's other same-class collision: build-a-website shares code modification with this family, so the dominant-house and dominant-element reading below repeats between the two pages exactly. The verdict distribution is where the two pages part company.

Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to writing code: reading code modification through a dominant house or a dominant element, 5 of the ten configurations; and, from this family's own canonical task, Inauspicious for 3 of them — the modal result in the distribution below.

Verdict distribution for writing code, ten registry configurations, fixed 1–7 scale order
Published valueVerdictConfigurations
1 of 7Inauspicious3 of the ten
2 of 7Unfavourable2 of the ten
3 of 7Guarded1 of the ten
4 of 7Neutral1 of the ten
5 of 7Favourable0 of the ten
6 of 7Auspicious1 of the ten
7 of 7Highly auspicious2 of the ten

Why writing code is classified as Code modification

The task both produces new source code and asserts, through its own tests, something about that code's correctness — squarely the code modification class, defined around exactly this combination of writing and verifying.

On the order form's task class field, select Code modification. That class's own risk-factor page is /tasks/code.

A written function's risk is that a passing test suite proves nothing about the paths it never exercised — a quiet, well-formatted failure rather than a loud one, which is the shape this class exists to track. 4 risk factors are tracked for code modification tasks generally; an assessment for writing code reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.

What this page does not claim about writing code

Which CSV dialect, which validation rules, and how many unit tests would be enough are all questions a real implementation answers and the canonical brief leaves open. The table below reads the brief as worded, not as any particular implementation would resolve it. The Registry does not test models and does not measure a success rate for writing code or any other task; what it does is compute, and the table at the top of this page is what it computed.

A buyer's own writing code task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.

Get Assessment

Related surfaces

← All capability questions · The AFR-1 method · Code modification task class

Related capability questions: Build a website · Open a link · Send an email

Questions about writing code

Can an AI model write code?
Across the ten registry configurations at their own registry-default settings, Inauspicious for 3 of them — the modal result for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Inauspicious and the newest returns Highly auspicious; the full ten-row table above lists the exact verdict for every configuration in between.
Why is writing code classified as Code modification rather than a different task class?
The task both produces new source code and asserts, through its own tests, something about that code's correctness — squarely the code modification class, defined around exactly this combination of writing and verifying.
Would a different programming language change which task class this family is assessed under?
No. Code modification covers writing, editing, refactoring, and testing code in any language; the canonical brief specifies Python because a brief needs one concrete wording, not because the classification is language-specific.