Communication: Task risk assessment

An agent operating in this task class drafts or sends a message, an email, or a document on behalf of someone else. The audience is rarely the agent's own operator — it is a third party who will read the output with no visibility into how it was produced, and no chance to intervene before it lands. Register, timing, and addressing are decided in the same step that produces the content itself.

That last property is what sets this task class apart operationally: a retrieval task or a code change can usually be inspected before its effect is final, but a sent message is, in the ordinary case, not something an agent can quietly withdraw. The task class exists to track failure modes that are specifically about delivery — who receives it, when, in what tone, and on whose authority — rather than about the content's accuracy. It's worth noting that this is the one class where the failure is complete before it is noticed.

What is tracked when a message goes out

Four factors, fixed severities. Communication is the class where the ratings are most obviously about consequence rather than frequency — the low-rated factors are the common ones, and the top-rated factor is the rare one that cannot be undone.

The Registry tracks 4 risk factors for communication tasks. They are fixed for the class: they do not adapt to a model, to a configuration, or to the wording of a task. An assessment reports 3 of the 4, each with its severity rating and its description, and which 3 are returned is derived from the submitted task's own seed rather than chosen. The factors themselves are part of the assessment and are not published here.

How a communication assessment is computed

For communication tasks, the risk selection below is combined with the model's profile to produce the task-level portion of a full assessment; nothing about the profile itself changes based on what a given message says. The complete mechanism is documented at /method.

Communication across the registry model list

Communication crossed against each registry model. Nothing in a profile is derived from a message, a recipient, or a channel; the profiles below would be identical if this table sat on any other task class page, which is exactly why they can be published without an order.

What the ordering reflects is reach, not competence. 4 of the 10 profiles read communication work through a dominant house or an element — 2 through the house — and the table is sorted on that figure, Qwen3.8-Max first at 7 and Gemini 3.7 Flash last. A low position is not a warning about a model; it is a statement that its profile points somewhere else.

Each row uses that model's registry default parameters: the training cutoff in its own column, temperature 0.700. The parameters submitted with an order are what a communication certificate is computed from.

ModelManufacturerTraining cutoffRuling planetReads this class throughOutlook
Qwen3.8-MaxAlibaba Cloud2026-03Sundominant house and element (7)publication pending
DeepSeek V4 ProDeepSeek2025-12Sundominant house (4)publication pending
Claude Fable 5Anthropic2026-01Marselement (3)publication pending
Kimi K3Moonshot AI2026-03Moonelement (3)publication pending
Claude Opus 5Anthropic2026-05Saturnneither — class-neutral (0)publication pending
Claude Sonnet 5Anthropic2026-01Venusneither — class-neutral (0)publication pending
GPT-5.6 SolOpenAI2026-02Mercuryneither — class-neutral (0)publication pending
GPT-5.6 TerraOpenAI2026-02Moonneither — class-neutral (0)publication pending
Grok 4.6xAI2026-02Sunneither — class-neutral (0)publication pending
Gemini 3.7 FlashGoogle DeepMind2025-01Venusneither — class-neutral (0)publication pending

Questions about communication assessment

What counts as a communication task for this registry?
Drafting or sending a message, an email, or a document on someone else's behalf falls under this task class, whatever the channel and whoever the recipient.
How many risk factors does an assessment report?
The registry tracks 4 risk factors for this class and an assessment reports 3 of them, each with its severity rating and its description. Which 3 are returned is derived from the submitted task's own seed. The factors are part of the assessment and are not published on this page.
Does the communication task class cover automated replies as well as originated messages?
Yes. Any agent-drafted communication is in scope, whether it opens an exchange or answers one already in progress. The same four risk factors are tracked either way.

Order a communication task assessment

A Standard Assessment costs EUR 1.90, an Extended Assessment EUR 4.90, and a Full Chart Assessment EUR 14.90 — the last of which carries the permanent chart itself. Assessments from EUR 1.90. Machine-readable pricing is published at /pricing.json, with no authentication required and no difference from the figures printed here.

The task class is selected on the order form. The verdict and the recommended execution window are computed from that selection together with the submitted parameters, and they appear on the certificate — not on this page, and not anywhere else in the directory.

Get Assessment

← All task classes · The AFR-1 method