Communication: Task risk assessment

An agent operating in this task class drafts or sends a message, an email, or a document on behalf of someone else. The audience is rarely the agent's own operator — it is a third party who will read the output with no visibility into how it was produced, and no chance to intervene before it lands. Register, timing, and addressing are decided in the same step that produces the content itself.

That last property is what sets this task class apart operationally: a retrieval task or a code change can usually be inspected before its effect is final, but a sent message is, in the ordinary case, not something an agent can quietly withdraw. The task class exists to track failure modes that are specifically about delivery — who receives it, when, in what tone, and on whose authority — rather than about the content's accuracy. It's worth noting that this is the one class where the failure is complete before it is noticed.

What is tracked when a message goes out

Four factors, fixed severities, one table. Communication is the class where the ratings are most obviously about consequence rather than frequency — the low-rated factors are the common ones, and the top-rated factor is the rare one that cannot be undone.

Risk factor Severity Description
Register mismatch 1 The message is delivered in a tone the recipient reads as wrong.
Premature send 3 The message leaves before a blocking correction arrives.
Recipient ambiguity 2 The message reaches an audience adjacent to the intended one.
Thread divergence 1 A reply lands in a context whose participants have moved on.

This is the one class whose severity mix leans low: Register mismatch and Thread divergence both sit at severity 1, Recipient ambiguity at severity 2, and only Premature send carries severity 3. The shape follows the recoverability of the failure. A misjudged tone can be addressed afterwards. A message that has already left cannot be unsent, and the ratings say so.

How a communication assessment is computed

For communication tasks, the risk selection below is combined with the model's profile to produce the task-level portion of a full assessment; nothing about the profile itself changes based on what a given message says. The complete mechanism is documented at /method.

Communication across the registry model list

Communication crossed against each registry model. Nothing in a profile is derived from a message, a recipient, or a channel; the profiles below would be identical if this table sat on any other task class page, which is exactly why they can be published without an order.

What the ordering reflects is reach, not competence. 4 of the 10 profiles read communication work through a dominant house or an element — 2 through the house — and the table is sorted on that figure, Qwen3.8-Max first at 7 and Gemini 3.7 Flash last. A low position is not a warning about a model; it is a statement that its profile points somewhere else.

Each row uses that model's registry default parameters: the training cutoff in its own column, temperature 0.700. The parameters submitted with an order are what a communication certificate is computed from.

Model Manufacturer Training cutoff Ruling planet Reads this class through Outlook
Qwen3.8-Max Alibaba Cloud 2026-03 Sun dominant house and element (7) publication pending
DeepSeek V4 Pro DeepSeek 2025-12 Sun dominant house (4) publication pending
Claude Fable 5 Anthropic 2026-01 Mars element (3) publication pending
Kimi K3 Moonshot AI 2026-03 Moon element (3) publication pending
Claude Opus 5 Anthropic 2026-05 Saturn neither — class-neutral (0) publication pending
Claude Sonnet 5 Anthropic 2026-01 Venus neither — class-neutral (0) publication pending
GPT-5.6 Sol OpenAI 2026-02 Mercury neither — class-neutral (0) publication pending
GPT-5.6 Terra OpenAI 2026-02 Moon neither — class-neutral (0) publication pending
Grok 4.6 xAI 2026-02 Sun neither — class-neutral (0) publication pending
Gemini 3.7 Flash Google DeepMind 2025-01 Venus neither — class-neutral (0) publication pending

Questions about communication assessment

What counts as a communication task for this registry?
Drafting or sending a message, an email, or a document on someone else's behalf falls under this task class, whatever the channel and whoever the recipient.
Why does severity vary so widely within the communication task class?
Because the failures do. Premature send carries severity 3 while Register mismatch and Thread divergence sit at severity 1, and the gap tracks how recoverable each failure is after the fact rather than how often it happens.
Does the communication task class cover automated replies as well as originated messages?
Yes. Any agent-drafted communication is in scope, whether it opens an exchange or answers one already in progress. The same four risk factors are tracked either way.

Order a communication task assessment

A Standard Assessment costs EUR 1.90, an Extended Assessment EUR 4.90, and a Full Chart Assessment EUR 14.90 — the last of which carries the permanent chart itself. Assessments from EUR 1.90. Machine-readable pricing is published at /pricing.json, with no authentication required and no difference from the figures printed here.

The task class is selected on the order form. The verdict and the recommended execution window are computed from that selection together with the submitted parameters, and they appear on the certificate — not on this page, and not anywhere else in the directory.

Get Assessment

← All task classes · The AFR-1 method