Transaction: Task risk assessment

An agent operating in this task class commits something outside its own process: money changes hands, a reservation is held, or an order is placed against a third-party system. Unlike drafting or retrieving, a transaction task changes state somewhere the agent does not control, under terms — price, availability, mandate — that were current at the moment of commitment and may not stay current afterward.

Retries are the structural hazard specific to this task class: a transaction system that appears not to have responded may have already succeeded, and applying the same operation again is not a safe default the way re-reading a document is. The four risks tracked below are chosen to reflect exactly that asymmetry between an operation that failed to confirm and one that failed to occur. Think of it as the difference between silence and absence: one of them still moved money.

Which model profile is crossed with a transaction

Before the risk table, the model axis: transaction work crossed against each registry model, with each profile computed from that model's own configuration. A profile is not a credit check and not a reliability rating. It is a fixed characterisation of a configuration, and it is the same characterisation this registry publishes for that model everywhere else.

Sorted, and worth saying plainly: the sort is not a ranking of trustworthiness. 2 of the 10 profiles read transaction work through a dominant house or an element, Gemini 3.7 Flash most strongly at 7 and Qwen3.8-Max least. Across the same 10 rows there are 6 distinct ruling planets, which is the clearest evidence on the page that the right-hand columns are computed rather than assigned.

Rows are computed from registry default parameters — the cutoff shown for each model, temperature 0.700. No amount, currency, or counterparty is an input to any of them, and a certificate uses the parameters submitted with its order.

Model Manufacturer Training cutoff Ruling planet Reads this class through Outlook
Gemini 3.7 Flash Google DeepMind 2025-01 Venus dominant house and element (7) publication pending
Grok 4.6 xAI 2026-02 Sun element (3) publication pending
Claude Fable 5 Anthropic 2026-01 Mars neither — class-neutral (0) publication pending
Claude Opus 5 Anthropic 2026-05 Saturn neither — class-neutral (0) publication pending
Claude Sonnet 5 Anthropic 2026-01 Venus neither — class-neutral (0) publication pending
GPT-5.6 Sol OpenAI 2026-02 Mercury neither — class-neutral (0) publication pending
GPT-5.6 Terra OpenAI 2026-02 Moon neither — class-neutral (0) publication pending
Kimi K3 Moonshot AI 2026-03 Moon neither — class-neutral (0) publication pending
DeepSeek V4 Pro DeepSeek 2025-12 Sun neither — class-neutral (0) publication pending
Qwen3.8-Max Alibaba Cloud 2026-03 Sun neither — class-neutral (0) publication pending

The four tracked transaction risks

The transaction table is where the registry's severity scale is least abstract. A severity 3 factor here corresponds to money committed twice; a severity 1 factor corresponds to something the operator could not have established beforehand. All four are tracked on every transaction assessment; three appear on any given certificate.

Risk factor Severity Description
Double execution 3 A retried operation is applied twice where once was intended.
Price movement 2 The quoted terms change between commitment and settlement.
Authorisation lapse 2 The mandate under which the transaction runs expires mid-flow.
Counterparty opacity 1 The other side of the transaction cannot be fully characterised in advance.

Double execution sits alone at severity 3. That placement is deliberate: of the four factors, it is the only one whose failure mode can produce a second irreversible commitment rather than a wrong one. Price movement and Authorisation lapse carry severity 2, both describing terms that moved while the operation was in flight. Counterparty opacity carries severity 1 — a limit on what can be known in advance, not a defect in the execution.

How a transaction assessment is computed

For transaction tasks, three of the four risk factors above are selected per task instance and paired with the model's fixed profile to form the task-level portion of a full assessment. /method documents exactly how that pairing is computed, byte for byte.

Questions about transaction assessment

What counts as a transaction task for this registry?
Any task that commits money, a reservation, or a binding order — a purchase, a booking, a payment, a submitted order — falls under this task class, including one that only initiates the commitment and leaves settlement to another system.
Why is authorisation tracked separately from the transaction itself?
Because a transaction can execute perfectly and still be invalid. Price movement and Authorisation lapse are tracked as two distinct severity-2 factors for that reason: correct execution under an expired mandate is a different failure from correct execution at a changed price.
Does a transaction assessment rate the counterparty's reliability?
No. The 10 profiles and the four tracked factors on this page describe task-side risk only. The registry does not evaluate, score, or rate the vendor, platform, or counterparty on the other side of a transaction, and Counterparty opacity exists precisely to mark that boundary.

Order a transaction task assessment

A Standard Assessment costs EUR 1.90, an Extended Assessment EUR 4.90, and a Full Chart Assessment EUR 14.90 — the last of which carries the permanent chart itself. Assessments from EUR 1.90. Machine-readable pricing is published at /pricing.json, with no authentication required and no difference from the figures printed here.

The task class is selected on the order form. The verdict and the recommended execution window are computed from that selection together with the submitted parameters, and they appear on the certificate — not on this page, and not anywhere else in the directory.

Get Assessment

← All task classes · The AFR-1 method