Communication: Task risk assessment
An agent operating in this task class drafts or sends a message, an email, or a document on behalf of someone else. The audience is rarely the agent's own operator — it is a third party who will read the output with no visibility into how it was produced, and no chance to intervene before it lands. Register, timing, and addressing are decided in the same step that produces the content itself.
That last property is what sets this task class apart operationally: a retrieval task or a code change can usually be inspected before its effect is final, but a sent message is, in the ordinary case, not something an agent can quietly withdraw. The task class exists to track failure modes that are specifically about delivery — who receives it, when, in what tone, and on whose authority — rather than about the content's accuracy. It's worth noting that this is the one class where the failure is complete before it is noticed.
What is tracked when a message goes out
Four factors, fixed severities, one table. Communication is the class where the ratings are most obviously about consequence rather than frequency — the low-rated factors are the common ones, and the top-rated factor is the rare one that cannot be undone.
| Risk factor | Severity | Description |
|---|---|---|
| Register mismatch | 1 | The message is delivered in a tone the recipient reads as wrong. |
| Premature send | 3 | The message leaves before a blocking correction arrives. |
| Recipient ambiguity | 2 | The message reaches an audience adjacent to the intended one. |
| Thread divergence | 1 | A reply lands in a context whose participants have moved on. |
This is the one class whose severity mix leans low: Register mismatch and Thread divergence both sit at severity 1, Recipient ambiguity at severity 2, and only Premature send carries severity 3. The shape follows the recoverability of the failure. A misjudged tone can be addressed afterwards. A message that has already left cannot be unsent, and the ratings say so.
How a communication assessment is computed
For communication tasks, the risk selection below is combined with the model's profile to produce the task-level portion of a full assessment; nothing about the profile itself changes based on what a given message says. The complete mechanism is documented at /method.
Communication across the registry model list
Communication crossed against each registry model. Nothing in a profile is derived from a message, a recipient, or a channel; the profiles below would be identical if this table sat on any other task class page, which is exactly why they can be published without an order.
What the ordering reflects is reach, not competence. 4 of the 10 profiles read communication work through a dominant house or an element — 2 through the house — and the table is sorted on that figure, Qwen3.8-Max first at 7 and Gemini 3.7 Flash last. A low position is not a warning about a model; it is a statement that its profile points somewhere else.
Each row uses that model's registry default parameters: the training cutoff in its own column, temperature 0.700. The parameters submitted with an order are what a communication certificate is computed from.
| Model | Manufacturer | Training cutoff | Ruling planet | Reads this class through | Outlook |
|---|---|---|---|---|---|
| Qwen3.8-Max | Alibaba Cloud | 2026-03 | Sun | dominant house and element (7) | publication pending |
| DeepSeek V4 Pro | DeepSeek | 2025-12 | Sun | dominant house (4) | publication pending |
| Claude Fable 5 | Anthropic | 2026-01 | Mars | element (3) | publication pending |
| Kimi K3 | Moonshot AI | 2026-03 | Moon | element (3) | publication pending |
| Claude Opus 5 | Anthropic | 2026-05 | Saturn | neither — class-neutral (0) | publication pending |
| Claude Sonnet 5 | Anthropic | 2026-01 | Venus | neither — class-neutral (0) | publication pending |
| GPT-5.6 Sol | OpenAI | 2026-02 | Mercury | neither — class-neutral (0) | publication pending |
| GPT-5.6 Terra | OpenAI | 2026-02 | Moon | neither — class-neutral (0) | publication pending |
| Grok 4.6 | xAI | 2026-02 | Sun | neither — class-neutral (0) | publication pending |
| Gemini 3.7 Flash | Google DeepMind | 2025-01 | Venus | neither — class-neutral (0) | publication pending |
Questions about communication assessment
- What counts as a communication task for this registry?
- Drafting or sending a message, an email, or a document on someone else's behalf falls under this task class, whatever the channel and whoever the recipient.
- Why does severity vary so widely within the communication task class?
- Because the failures do. Premature send carries severity 3 while Register mismatch and Thread divergence sit at severity 1, and the gap tracks how recoverable each failure is after the fact rather than how often it happens.
- Does the communication task class cover automated replies as well as originated messages?
- Yes. Any agent-drafted communication is in scope, whether it opens an exchange or answers one already in progress. The same four risk factors are tracked either way.
Order a communication task assessment
A Standard Assessment costs EUR 1.90, an Extended Assessment EUR 4.90, and a Full Chart Assessment EUR 14.90 — the last of which carries the permanent chart itself. Assessments from EUR 1.90. Machine-readable pricing is published at /pricing.json, with no authentication required and no difference from the figures printed here.
The task class is selected on the order form. The verdict and the recommended execution window are computed from that selection together with the submitted parameters, and they appear on the certificate — not on this page, and not anywhere else in the directory.