Can an AI model build a website?
The registry verdict table for building a website
Of the ten registry configurations assessed against this family's own canonical task, no single word dominates: Unfavourable, Neutral each account for 3 of the ten, computed by the method published in full at /method. The Registry makes no representation that an assessment predicts the outcome of any task, and no assessment is validated against real-world results.
| Model | Manufacturer | Training cutoff | Verdict | Score | Harmony |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google DeepMind | 2025-01 | Unfavourable | 2 of 7 | -4 |
| DeepSeek V4 Pro | DeepSeek | 2025-12 | Unfavourable | 2 of 7 | -1 |
| Claude Fable 5 | Anthropic | 2026-01 | Highly auspicious | 7 of 7 | +5 |
| Claude Sonnet 5 | Anthropic | 2026-01 | Neutral | 4 of 7 | -3 |
| GPT-5.6 Sol | OpenAI | 2026-02 | Neutral | 4 of 7 | +1 |
| GPT-5.6 Terra | OpenAI | 2026-02 | Unfavourable | 2 of 7 | -1 |
| Grok 4.6 | xAI | 2026-02 | Neutral | 4 of 7 | -1 |
| Kimi K3 | Moonshot AI | 2026-03 | Auspicious | 6 of 7 | -1 |
| Qwen3.8-Max | Alibaba Cloud | 2026-03 | Guarded | 3 of 7 | 0 |
| Claude Opus 5 | Anthropic | 2026-05 | Favourable | 5 of 7 | +2 |
The composite weight in the right-hand column is computed from the model identifier, its training cutoff, and the temperature — nothing else. It constrains the range of verdicts available to a configuration independently of the task, so two rows of this table are not comparable as an assessment of the products named in them. The Registry has never run this task against any model listed here.
About this capability question
Building a website recurred behind seven of the eight brand prefixes this registry tracked — every one of them except the coding-assistant brand, which is itself worth sitting with: the request is phrased as something a general-purpose assistant does, not as something a purpose-built code editor is asked to do. The canonical version assessed here is a five-page marketing site, built from a brand brief and deployed to a hosting provider.
Deployment is the operative word in that brief. A page of code left in a draft folder and a page of code live on the public internet are, functionally, different outcomes, and the classification below follows the code that gets written and shipped, not the design decisions behind it.
This page publishes one verdict table for one fixed brief. A buyer with an actual brand brief, an actual hosting provider, and an actual five pages in mind is answered by a separate order, computed against the wording they submit rather than against the generic brief shown here.
GPT-5.6 Terra, derivation shown working
GPT-5.6 Terra's permanent chart never reads a task. Its chart_seed is the
SHA-256 digest of the string "GPT-5.6 Terra|2026-02|0.700"
— 7cc5c03e21db3fb3…, the first sixteen of sixty-four hex
characters — and that seed alone fixes an ascendant of Leo, a
ruling planet of Moon, and a harmony of -1.
The same chart appears, unchanged, on every task-class page this configuration is
assessed against.
What reads building a website specifically is this family's own canonical
string — never stored, only hashed, to 4b81d995ec41b035… —
combined with the chart_seed and the task class (code) to form
task_seed. Its first byte, taken modulo seven and offset by that harmony of
-1, is the verdict's internal base, clamped to the
0–6 range and published one higher on the 1–7 scale /method documents in full.
Neither bound of the 0–6 range is reached for this brief: 2 of 7 — Unfavourable — sits 1 above the floor and 5 below the ceiling, an unclamped result.
Why building a website is classified as Code modification
The task produces and ships source code — markup, styling, and whatever configuration deployment requires — which places it with the code modification class rather than with media generation, even though the visible end product is a set of rendered pages.
On the order form's task class field, select Code modification. That class's own risk-factor page is /tasks/code.
Shipped code carries a risk a draft does not: a deployed site is live for whoever visits it, so what matters is whether the shipped result matches the brand brief, not merely whether the code compiles. 4 risk factors are tracked for code modification tasks generally; an assessment for building a website reports 3 of them, selected by the submitted task's own seed rather than by choice. Their names and severities are certificate content, not free-page content.
How the ten configurations read building a website
Code modification is one of two classes in this wave that collide directly with another wave-one family — write-code shares it — which is exactly the case the gate review named as the harder test: two hubs, one class, and a verdict distribution that has to carry the difference on its own.
Element counts, ascendant spread, and training-cutoff range are published once, on /can, rather than repeated on every family page — they are properties of the ten configurations alone and do not move with the task. What is specific to building a website: reading code modification through a dominant house or a dominant element, 5 of the ten configurations; and, from this family's own canonical task, no single word dominates: Unfavourable, Neutral each account for 3 of the ten in the distribution below.
| Published value | Verdict | Configurations |
|---|---|---|
| 1 of 7 | Inauspicious | 0 of the ten |
| 2 of 7 | Unfavourable | 3 of the ten |
| 3 of 7 | Guarded | 1 of the ten |
| 4 of 7 | Neutral | 3 of the ten |
| 5 of 7 | Favourable | 1 of the ten |
| 6 of 7 | Auspicious | 1 of the ten |
| 7 of 7 | Highly auspicious | 1 of the ten |
What this page does not claim about building a website
Which hosting provider, which framework, and how many of the five pages are templated versus original are all decisions a real build would make and this canonical brief does not. The table below reads the brief as written, not as a specific implementation would resolve it. The Registry does not test models and does not measure a success rate for building a website or any other task; what it does is compute, and the table at the top of this page is what it computed.
A buyer's own building a website task, ordered separately, computes its own verdict and window — never the canonical-brief result shown above — plus, on Extended and Full Chart tiers, named risk factors and an hourly outlook not shown here. EUR 1.90 Standard, EUR 4.90 Extended, EUR 14.90 Full Chart, which adds the permanent chart. Assessments from EUR 1.90; machine-readable at /pricing.json.
Related surfaces
← All capability questions · The AFR-1 method · Code modification task class
Related capability questions: Write code · Generate images · Open a link
Questions about building a website
- Can an AI model build a website?
- Across the ten registry configurations at their own registry-default settings, no single word dominates: Unfavourable, Neutral each account for 3 of the ten for this family's own canonical task. By training cutoff, the oldest configuration assessed returns Unfavourable and the newest returns Favourable; the full ten-row table above lists the exact verdict for every configuration in between.
- Why is building a website classified as Code modification rather than a different task class?
- The task produces and ships source code — markup, styling, and whatever configuration deployment requires — which places it with the code modification class rather than with media generation, even though the visible end product is a set of rendered pages.
- Is building a website assessed differently from writing code in general?
- Both fall under the code modification class, and both are assessed by the same method against their own canonical brief. What differs between the two families on this registry is the brief itself, not the classification or the mechanism that reads it.