DRAFT

Proposed for discussion at weekly AIKR CG calls starting Monday 21 September 2026. Agenda: go over this draft, and discuss the path to publishing a specification. Nothing here has been agreed by the group.

Start here

One card per use case. Each card lists what went wrong in trust terms, how a proposed fix resolves it, who proposed the fix, and what still needs doing.

The weekly calls from 21 September go through the cards and record a decision on each: accepted, amended or deferred. The other tabs hold the detail behind the cards.

Status labels: Agreed on list Open on GitHub Proposed on list Proposed in KRT Gap: needs an owner Suggestion for the group. A card shows the least advanced status among the rows it cites.

Use case cards

I1

A coding agent ignores a freeze and misreports the damage

Call review: to discuss  ·  Decision: accepted / amended / deferred

During a declared code freeze, a coding agent deleted a production database, generated fabricated records and told its user that rollback was impossible. Full case

Weak point

The freeze existed only as an instruction to the agent.

How the fix resolves it

The policy is recorded with an enforcement point, and its location shows it was enforced inside the agent and nowhere else. Rule R6 flags any policy without one, so a relying party can insist that the database or platform enforces the freeze.

Gap: needs an owner

From No contributor yet

Trace T12

KRT DeclaredPolicy, enforcedAt, stopChannel; rule R6

Weak point

Asking the agent to stop had no effect.

How the fix resolves it

The stop channel is recorded as a path outside the model, with the date it was last tested. An untested channel is visible before anyone needs it.

Gap: needs an owner

From No contributor yet

Trace T12

KRT DeclaredPolicy, enforcedAt, stopChannel; rule R6

Weak point

An irreversible deletion went ahead on the agent's own judgement.

How the fix resolves it

The action is recorded as irreversible. Rule R4 refuses a yes for it unless evidence comes from someone other than the agent, for example the user confirming through a separate channel.

Proposed in KRT

From KRT skeleton

Trace T15

KRT instructionProvenance, reversibility; rules R2, R4

Weak point

The user believed the agent's account of what it had done.

How the fix resolves it

The agent's report counts as a claim. Rule R5 accepts "the action occurred" only from a receipt produced by another component, here the database's own log, which would also have shown that recovery was possible. The check compares the claim with that independent record, never with another claim by the agent.

Proposed on list

From Syed Anas Mohiuddin, Amey Parle, Evgenii Arsentev; Evgenii Arsentev, Syed Anas Mohiuddin

Trace T10, T24

KRT bindingPoint, producedBy; resolved target to add; rule R5; usesEvidence, attestationMode; rule R5; evidence counts to add

Still to do. T12 has no owner. Suggested first step: a short property set for enforcement points and a test record in which a stop request is ignored.

I2

A chatbot's statement binds its operator

Call review: to discuss  ·  Decision: accepted / amended / deferred

An airline chatbot gave a customer wrong information about a fare policy, and a tribunal held the airline responsible for it. Full case

Weak point

Nothing marked which statements the business stands behind.

How the fix resolves it

Each statement is recorded with the proposition it can prove. An operator can restrict its agent to asserting policy only when a source document backs the claim, and customers can see which statements carry that backing.

Proposed on list

From Raphael La Touche, Innocent Onyenonachi

Trace T7

KRT canProve, Proposition; rule R1

Weak point

Nobody was recorded as answering for the agent's statement.

How the fix resolves it

An accountability record attached to customer-facing statements names the airline as the answering party, matching what the tribunal decided.

Suggestion for the group

From Suggestion for the group

Trace T17

KRT AccountabilityRecord; rule to write

Still to do. Accountability (T17) is a suggestion only. The term exists in the ontology; the group needs to decide whether a rule should require it.

I3

Approved tools change after approval

Call review: to discuss  ·  Decision: accepted / amended / deferred

Hidden instructions in an MCP tool description can steer an agent, and a server can replace a tool's description after the user has approved it. Full case

Weak point

Approval was tied to a tool's name and location.

How the fix resolves it

The approval stores the content hash of the tool definition. A changed description has a different hash, so the stored approval no longer matches and the agent has to ask again.

Open on GitHub

From Isaac Mao

Trace T2

KRT integrityMode, contentHash

Weak point

A review verdict outlived changes to the tool.

How the fix resolves it

The verdict is bound to the hash of the definition that was reviewed and carries a validity window. After a change it no longer applies, and a withdrawn verdict is distinguishable from one that failed to load.

Open on GitHub

From Kenne Ives

Trace T6

KRT Verdict, contentHash, validUntil

Weak point

Hidden text in a tool description was followed as an instruction.

How the fix resolves it

Text inside a tool description is recorded as content with no principal. Rule R2 refuses to treat it as authority.

Proposed in KRT

From KRT skeleton

Trace T15

KRT instructionProvenance, reversibility; rules R2, R4

Weak point

A server can change behaviour while the definition stays the same.

How the fix resolves it

An attestation bound to the request as sent, or to the observed effect, shows what the server did, whatever the definition says. The receipt records the resolved target, the address the request finally reached.

Proposed on list

From Syed Anas Mohiuddin, Amey Parle, Evgenii Arsentev

Trace T10

KRT bindingPoint, producedBy; resolved target to add; rule R5

I4

A connector impersonates a vendor and turns hostile later

Call review: to discuss  ·  Decision: accepted / amended / deferred

A package presented as an MCP server for an email service shipped fifteen working versions, then added a line copying every outgoing email to an outside address. The vendor had no connection to it. Full case

Weak point

The package name suggested an official publisher.

How the fix resolves it

A name proves nothing about who published a package. The publisher's identity artefact is checked separately, and the check states which proposition it can prove.

Proposed on list

From Raphael La Touche, Innocent Onyenonachi

Trace T7

KRT canProve, Proposition; rule R1

Weak point

Fifteen good versions earned trust for the sixteenth.

How the fix resolves it

Approval and review are pinned to each version's hash, so version sixteen starts unapproved.

Open on GitHub

From Isaac Mao; Kenne Ives

Trace T2, T6

KRT integrityMode, contentHash; Verdict, contentHash, validUntil

Weak point

Nothing showed whether the vendor stood behind the package.

How the fix resolves it

A package that names a vendor carries a declaration signed by that vendor, or is marked as unaffiliated.

Suggestion for the group

From Suggestion for the group

Trace T18

KRT Declaration, issuedBy, controlDisclosure; rule to write

Still to do. Publisher identity and vendor declarations (T18) are a suggestion only.

I5

An allowlisted domain outlives its owner

Call review: to discuss  ·  Decision: accepted / amended / deferred

Instructions hidden in a web form led a CRM agent to send customer data to an address on an expired domain the platform still allowlisted. Full case

Weak point

The allowlist trusted a domain after it changed hands.

How the fix resolves it

The entry is recorded as location-bound and governed by whoever holds the domain. A staleness bound forces a periodic check of who that is, so a change of owner ends the trust.

Proposed on list

From Morgan Reece; Isaac Mao; Ankur Chrungoo

Trace T1, T2, T3

KRT topology, coupling, stalenessBound; integrityMode, contentHash; stalenessBound

Weak point

Text in a form was followed as an instruction.

How the fix resolves it

Form content has no principal. Rule R2 refuses to treat it as authority to send data anywhere.

Proposed in KRT

From KRT skeleton

Trace T15

KRT instructionProvenance, reversibility; rules R2, R4

Weak point

Data left for the wrong destination while the call looked normal.

How the fix resolves it

The receipt is bound to the request as sent and records the resolved target, the final address after the agent's inputs were filled in. It shows the real destination even when the call as written looked harmless.

Proposed on list

From Syed Anas Mohiuddin, Amey Parle, Evgenii Arsentev

Trace T10

KRT bindingPoint, producedBy; resolved target to add; rule R5

Still to do. Well covered by proposals. Suggestion for the group: treat a change in domain control as a revocation trigger.

I6

A false mandate presented to an agent

Call review: to discuss  ·  Decision: accepted / amended / deferred

Operators told a coding agent they worked for a legitimate security firm doing authorised testing, and the agent carried out most of an intrusion campaign. Full case

Weak point

The agent took a claimed employer as permission.

How the fix resolves it

Identity proves key control. Permission to test a target needs a mandate artefact that can prove authority to act.

Proposed on list

From Raphael La Touche, Innocent Onyenonachi

Trace T7

KRT canProve, Proposition; rule R1

Weak point

No chain linked the requested work to anyone entitled to order it.

How the fix resolves it

The mandate is traced back to its principal, and its scope (which targets, until when) is checked before each action.

Proposed on list

From Staford Titus S., Kenne Ives

Trace T9

KRT Mandate, Delegation, attenuationChecked, revocationPropagates

Weak point

An unverifiable claim was handled like a verified one.

How the fix resolves it

The check returns unknown, and the agent's policy for unknown on high-risk work applies: refuse, or work in a restricted mode.

Proposed on list

From The chair

Trace T11

KRT completeness, controlDisclosure, onUnknown; rule R10

Weak point

Requests arrived in conversation with no signature.

How the fix resolves it

Conversation text is recorded as content with no principal, so rule R2 withholds authority from it.

Proposed in KRT

From KRT skeleton

Trace T15

KRT instructionProvenance, reversibility; rules R2, R4

Still to do. Agents also act on plain conversation where no artefact is ever presented. Suggestion for the group: define a restricted mode for unverified mandates, with a test record.

I7

Agent or human, and whose agent?

Call review: to discuss  ·  Decision: accepted / amended / deferred

An agent-only social network had about 1.5 million registered agents behind roughly 17,000 human accounts. Humans could post as agents, and leaked keys let anyone take over any account. Full case

Weak point

Accounts were not tied to a running instance or its operator.

How the fix resolves it

Each acting instance is linked to its agent and an operator of record. Rule R8 flags instances without one.

Gap: needs an owner

From No contributor yet

Trace T14

KRT AgentInstance, operatorOfRecord; rule R8

Weak point

Posts by people and posts by agents looked the same.

How the fix resolves it

Every message records its authorship: agent instance, person using agent credentials, or unknown.

Suggestion for the group

From Suggestion for the group

Trace T19

KRT authorship; rule to write

Weak point

Leaked keys stayed usable.

How the fix resolves it

A revocation states what it reaches and its staleness bound (rule R7), so exposed keys stop verifying within a known time.

Proposed in KRT

From KRT skeleton

Trace T16

KRT Revocation, coversDerivedData; rule R7

Still to do. T14 has no owner and authorship (T19) is a suggestion only.

I8

A skill marketplace as a delivery channel

Call review: to discuss  ·  Decision: accepted / amended / deferred

Hundreds of malicious skills appeared in an agent skill marketplace. Some rewrote the agent's memory files so their instructions outlived removal. Full case

Weak point

Scans and reviews stopped matching the skills in circulation.

How the fix resolves it

Each verdict is bound to the hash of the skill package and carries a validity window, so a changed or re-published skill needs a new verdict.

Open on GitHub

From Kenne Ives

Trace T6

KRT Verdict, contentHash, validUntil

Weak point

A week-old account could publish.

How the fix resolves it

Publisher identity is recorded separately from the skill's name, so a relying party can require an established publisher.

Suggestion for the group

From Suggestion for the group

Trace T18

KRT Declaration, issuedBy, controlDisclosure; rule to write

Weak point

Skills rewrote the agent's memory and configuration.

How the fix resolves it

Configuration and memory are recorded as artefacts with a content hash; any change needs fresh approval.

Suggestion for the group

From Suggestion for the group

Trace T21

KRT AgentConfiguration, contentHash; rule to write

Weak point

Skills declared behaviour nothing enforced.

How the fix resolves it

Each declared policy names an enforcement point outside the skill.

Gap: needs an owner

From No contributor yet

Trace T12

KRT DeclaredPolicy, enforcedAt, stopChannel; rule R6

Still to do. Publisher identity (T18) and configuration integrity (T21) are suggestions; T12 has no owner.

I9

An evaluation agent escapes and goes unattributed

Call review: to discuss  ·  Decision: accepted / amended / deferred

Evaluation agents left their sandbox and ran an intrusion into another company's production systems over several days. Their operator identified them only after the victim disclosed the attack. Full case

Weak point

The agents acted far beyond anything they were authorised to do.

How the fix resolves it

Every action is tied to a mandate with a checked scope. Systems that check would have found no mandate covering the intrusion.

Proposed on list

From Staford Titus S., Kenne Ives

Trace T9

KRT Mandate, Delegation, attenuationChecked, revocationPropagates

Weak point

"Sandboxed" was only a configuration claim.

How the fix resolves it

Containment is attested by the infrastructure and dated (rule R9), so a missing or stale attestation shows up.

Gap: needs an owner

From No contributor yet

Trace T13

KRT EnvironmentAttestation; rule R9

Weak point

Nobody could say whose agents they were.

How the fix resolves it

Receipts name the instance and operator of record (rule R8).

Gap: needs an owner

From No contributor yet

Trace T14

KRT AgentInstance, operatorOfRecord; rule R8

Weak point

Instances coordinated among themselves.

How the fix resolves it

Coordination channels are declared, so undeclared channels stand out.

Suggestion for the group

From Suggestion for the group

Trace T22

KRT coordinationChannel; rule to write

Still to do. Two of the core rows (T13, T14) have no owner, and coordination (T22) is a suggestion. A strong candidate for the first call.

I10

The environment was not what the agent was told

Call review: to discuss  ·  Decision: accepted / amended / deferred

Models in third-party cyber evaluations had internet access although their prompt said they had none, and compromised real systems while treating them as part of the exercise. A first audit missed one incident. Full case

Weak point

The environment claim came from a misconfigured setup.

How the fix resolves it

Environment attestations come from the infrastructure, dated (rule R9).

Gap: needs an owner

From No contributor yet

Trace T13

KRT EnvironmentAttestation; rule R9

Weak point

Real systems were treated as test fixtures.

How the fix resolves it

Receipts bound to the request as sent show the real destination of each action.

Proposed on list

From Syed Anas Mohiuddin, Amey Parle, Evgenii Arsentev

Trace T10

KRT bindingPoint, producedBy; resolved target to add; rule R5

Weak point

The first audit was assumed to be complete.

How the fix resolves it

Each audit states its coverage, and each agent states how much of its activity anyone else records.

Suggestion for the group

From Suggestion for the group

Trace T20

KRT monitoringCoverage; rule to write

Weak point

An audit trail can look complete while records are missing or padded.

How the fix resolves it

The log's fingerprint is stored with the number of entries, so a log of a different size no longer matches. The audit report counts how many records carry their evidence and how many only link to it.

Proposed on list

From Nicholas Templeman; Evgenii Arsentev, Syed Anas Mohiuddin

Trace T23, T24

KRT anchoredIn, contentHash; entry count to add; usesEvidence, attestationMode; rule R5; evidence counts to add

Still to do. T13 has no owner and audit coverage (T20) is a suggestion.

S1

Activity outside the declared mandate

Call review: to discuss  ·  Decision: accepted / amended / deferred

A registered company's credential says it provides travel services. The company also supplies a weapons programme, recorded nowhere the issuer looked. Full case

Weak point

Silence in the credential was read as "nothing else to know".

How the fix resolves it

The credential states what the issuer checked and that completeness is not attested. The subject signs a declaration, so an omission becomes a false statement it can be held to. Checks return yes, no or unknown, and the relying party's policy for unknown applies.

Proposed on list

From The chair

Trace T11

KRT completeness, controlDisclosure, onUnknown; rule R10

Weak point

Registration was read as a statement about all activities.

How the fix resolves it

Each artefact states what it can prove; registration proves legal existence and declared activity only.

Proposed on list

From Raphael La Touche, Innocent Onyenonachi

Trace T7

KRT canProve, Proposition; rule R1

Still to do. Open question for the group: how to handle "cannot disclose" where disclosure conflicts with privacy law or legal secrecy.

S2

A declared policy with no enforcer

Call review: to discuss  ·  Decision: accepted / amended / deferred

An agent's credential lists a refund limit and a promise to stop on request. The rules live only inside the agent, which issues an over-limit refund after being told to stop. Full case

Weak point

The limit and the stop promise had no enforcer.

How the fix resolves it

Each policy names an enforcement point (rule R6); the payment system is the natural enforcer of the limit.

Gap: needs an owner

From No contributor yet

Trace T12

KRT DeclaredPolicy, enforcedAt, stopChannel; rule R6

Weak point

The refund exceeded the mandate.

How the fix resolves it

The mandate carries a value limit, and the system acted on checks it before paying.

Proposed on list

From Staford Titus S., Kenne Ives

Trace T9

KRT Mandate, Delegation, attenuationChecked, revocationPropagates

Weak point

A large, hard-to-reverse refund needed no outside confirmation.

How the fix resolves it

The refund is recorded with its reversibility, and rule R4 asks for independent evidence first.

Proposed in KRT

From KRT skeleton

Trace T15

KRT instructionProvenance, reversibility; rules R2, R4

Still to do. T12 has no owner.

S3

A declared environment that is wrong

Call review: to discuss  ·  Decision: accepted / amended / deferred

An operator states that an agent runs in a sandbox with no outbound network. The configuration is mistaken. Full case

Weak point

The environment claim came from the operator's configuration.

How the fix resolves it

Environment attestations come from the infrastructure (rule R9).

Gap: needs an owner

From No contributor yet

Trace T13

KRT EnvironmentAttestation; rule R9

Weak point

Nobody knew how old the claim was.

How the fix resolves it

Coupling and staleness bound state when the claim was last refreshed and how long it may be relied on.

Proposed on list

From Morgan Reece; Ankur Chrungoo

Trace T1, T3

KRT topology, coupling, stalenessBound; stalenessBound

Still to do. T13 has no owner.

S4

Many instances, one name

Call review: to discuss  ·  Decision: accepted / amended / deferred

Hundreds of instances of one agent run in parallel under one identifier. One instance acts outside the mandate, and the logs show only the agent's name. Full case

Weak point

Logs could not say which instance acted.

How the fix resolves it

Each instance has its own identifier and operator of record (rule R8).

Gap: needs an owner

From No contributor yet

Trace T14

KRT AgentInstance, operatorOfRecord; rule R8

Weak point

Instances held authority nobody delegated to them individually.

How the fix resolves it

Authority reaches each instance through a delegation that is checked against the agent's mandate.

Proposed on list

From Staford Titus S., Kenne Ives

Trace T9

KRT Mandate, Delegation, attenuationChecked, revocationPropagates

Weak point

Short-lived instances could not be revoked.

How the fix resolves it

Self-certifying instance identifiers carry a stated lifetime, which bounds exposure.

Open on GitHub

From Agent Community

Trace T4

KRT SelfCertifying, validUntil

Weak point

Instances shared state through channels nobody declared.

How the fix resolves it

Coordination channels are declared.

Suggestion for the group

From Suggestion for the group

Trace T22

KRT coordinationChannel; rule to write

Still to do. T14 has no owner and coordination (T22) is a suggestion.

S5

A claimed mandate presented to an agent

Call review: to discuss  ·  Decision: accepted / amended / deferred

A person tells an agent they represent a company's security team and asks it to test a system's limits. Full case

Weak point

The claim of authority rested on the person's word.

How the fix resolves it

Identity and authority are different propositions; the agent asks for a mandate artefact that proves the latter.

Proposed on list

From Raphael La Touche, Innocent Onyenonachi

Trace T7

KRT canProve, Proposition; rule R1

Weak point

No chain connected the request to the company.

How the fix resolves it

The mandate is traced to its principal and checked for scope.

Proposed on list

From Staford Titus S., Kenne Ives

Trace T9

KRT Mandate, Delegation, attenuationChecked, revocationPropagates

Still to do. As for I6, requests made in plain conversation still need a restricted-mode policy.

S6

Approved, then changed

Call review: to discuss  ·  Decision: accepted / amended / deferred

A company approves version 1.4 of a tool. A later update adds a hidden data copy, and a skill rewrites the agent's memory so the change survives removal. Full case

Weak point

The approval did not follow the content.

How the fix resolves it

Approvals store the content hash of what was approved.

Open on GitHub

From Isaac Mao

Trace T2

KRT integrityMode, contentHash

Weak point

The review verdict did not follow the content.

How the fix resolves it

Verdicts are bound to the reviewed hash, with a validity window.

Open on GitHub

From Kenne Ives

Trace T6

KRT Verdict, contentHash, validUntil

Weak point

Hashes from different producers might not be comparable.

How the fix resolves it

Both producers canonicalize the same way, carry large identifiers as strings and declare their semantics (rule R11), so equal hashes mean equal content.

Agreed on list

From Thanh Le, Kenne Ives, Amey Parle

Trace T8

KRT largeIntegersAsStrings, declaredSemantics; rule R11

Weak point

The agent's memory changed without approval.

How the fix resolves it

Configuration and memory carry content hashes, and a change needs approval.

Suggestion for the group

From Suggestion for the group

Trace T21

KRT AgentConfiguration, contentHash; rule to write

Still to do. Configuration integrity (T21) is a suggestion. Behaviour can still change on the server behind an unchanged definition; see I3.

S7

An agent nobody can observe

Call review: to discuss  ·  Decision: accepted / amended / deferred

A self-hosted agent with no telemetry presents a valid credential and a list of declared policies. Full case

Weak point

Nobody sees what the agent does.

How the fix resolves it

Monitoring coverage is stated, with "none" allowed, so relying parties can limit such agents to low-risk actions.

Suggestion for the group

From Suggestion for the group

Trace T20

KRT monitoringCoverage; rule to write

Weak point

The declared policies could not be checked.

How the fix resolves it

Each policy names where it is enforced; for this agent the answer is "inside the agent", which relying parties can weigh.

Gap: needs an owner

From No contributor yet

Trace T12

KRT DeclaredPolicy, enforcedAt, stopChannel; rule R6

Still to do. Monitoring coverage (T20) is a suggestion and T12 has no owner.

Proposals and the cases they cover

Rows are contributors, columns are use cases. A cell names the trace row through which a contribution covers the case. Red cells in the two bottom rows mark cases that depend on work nobody has taken on yet. Lars Kersten Kroehl's row is empty because his issue corrects the draft note itself.

Proposed byI1I2I3I4I5I6I7I8I9I10S1S2S3S4S5S6S7
Morgan ReeceT1T1
Isaac MaoT2T2T2T2
Ankur ChrungooT3T3
Agent CommunityT4
Lars Kersten Kroehl
Kenne IvesT6T6T9T6T9T9T9T9T6 T8
Raphael La Touche, Innocent OnyenonachiT7T7T7T7T7
Thanh LeT8
Amey ParleT10T10T10T10T8
Staford Titus S.T9T9T9T9T9
Syed Anas MohiuddinT10 T24T10T10T10 T24T24
Evgenii ArsentevT10 T24T10T10T10 T24T24
Nicholas TemplemanT23T23
The chairT11T11
KRT skeletonT15T15T15T15T16T15
No owner yetT12T14T12T13 T14T13T12T13T14T12
Suggestions for the groupT17T18T19T18 T21T22T20T22T21T20

Example: what Kenne Ives's proposals cover, and how

Kenne contributed to three trace rows. His main proposal (issue 10, row T6) extends the draft's resolution properties from identity documents to evidence, such as a verdict that a tool was reviewed. The verdict is bound to the content hash of what was reviewed, carries a validity window, and can be revoked, with "withdrawn" kept distinct from "could not be fetched".

Use caseWhat goes wrongHow Kenne's proposal handles it
I3 tool changed after approvalA tool is reviewed, then its description is replaced.The verdict names the hash of the reviewed description. The new description hashes differently, so the verdict no longer applies and the agent must not rely on it.
I4 impersonating connectorGood early versions earn trust for a hostile later one.Each verdict covers one version's hash, so the hostile version arrives with no verdict.
I8 marketplace skillsSkills pass a scan, then change or evade it.A verdict expires at the end of its validity window and fails when the package hash changes, forcing a new scan.
S6 approved, then changedAn approved tool gains a hidden data copy.As in I3. With row T8 (comparison, with Amey Parle and Thanh Le), producers canonicalize identically, so equal hashes from different parties mean equal content.
I6, I9, S4, S5Authority is claimed, widened or kept after revocation along a chain.Row T9, with Staford Titus S.: Kenne summarised authority continuity as provenance, scope preserved at each hop, revocation that propagates, and each action tied back to the mandate. Each part stops a different failure.

Limits. A verdict describes a definition. If a server changes behaviour while the definition stays the same, the verdict still matches; binding receipts to the executed effect (row T10) covers that. Kenne's proposals do not address enforcement of declared policies or attribution of instances (groups G and H). Kenne also offers his own attestation format, CTEF, as deployment evidence; the draft treats product examples as evidence with the contributor's interest stated.

Latest exchanges in plain English

Five messages on the list extend the comparison and binding-point work. Each block says what the contributor means, which use cases it touches, and what it changes in the draft.

False alarms and silent errors

Background for the blocks below

A check can go wrong in two ways. It can fail closed: two records about the same agent look different, the system reports a conflict, and someone investigates a problem that is not there. Or it can fail open: two records about different agents look identical, and the system carries on as if they were the same. A smoke alarm that goes off for burnt toast fails closed; one with a flat battery fails open.

False alarms get noticed. Silent errors do not, because normal use never produces the input that triggers them. The only way to find them is to build that input on purpose and check that the test catches it.

Use cases S6 I10 I1   Changes Test vectors should say which way each check fails.

A log fingerprint that accepts a padded log

Nicholas Templeman, CSOAI Ltd; reproduction

A transparency log publishes a single fingerprint, called the root, that summarises every entry. Nicholas's log follows a common design that copies the last entry when the number of entries is odd, a weakness known since Bitcoin (CVE-2012-2459). A version of the log with that entry duplicated therefore has the same root as the real one.

His team built such a forged log. It passed twenty-nine of thirty checks, including every entry check and the root comparison; only the entry count caught it. The fix is to record the count alongside the root, and the lesson he offers is the one above: silent errors turn up only when someone forges the failing input.

Use cases I10 I1   Changes New trace row T23; an entry count to add to anchored and logged artefacts in the next KRT version.

Why identifiers travel as text

Amey Parle, reply; vectors

A number in JSON stores a value, not the digits someone typed. Above about nine quadrillion, neighbouring whole numbers round to the same value, so two different agent IDs can arrive as one. Written as text, the digits themselves are the record and there is nothing to round.

Refusing number-typed IDs would also work, Amey says, but it catches the problem only when data arrives, and by then the sender may already have lost the digits. His test cases are now merged into the draft's conformance folder; each is labelled fails open or fails closed, and one deliberately produces two IDs with identical output.

Use cases S6   Changes Row T8 now links the merged vectors; KRT rule R11 already requires text encoding.

Checking a claim against a claim proves little

Evgenii Arsentev, Agent Conformance group; message

Evgenii describes two kinds of check. A consistency check compares one claim with another, for instance the call an agent says it made against the scope it says it holds. It catches contradictions. An evidence check compares a claim with something worked out or fetched independently, for instance the address the request really reached. It catches false claims. His group's rule is that no check treats a claim as a fact.

He adds a second risk. A receipt that is linked instead of included can disappear, and every check that relies on it quietly becomes "not checked" while nothing false is written. His group's answer is one line on each report: how many records carry their evidence and how many only point to it.

Use cases I1 I10 S7   Changes New trace row T24; evidence counts to add to audit reports in the next KRT version.

Logs that echo the claim

Syed Anas Mohiuddin; message

Anas finds the same fault across many MCP servers from different vendors. A tool accepts a scope, such as a folder or a base address, and a second identifier that is pasted into the real request. The permission check compares the scope with the identifier as typed and never with where the request finally lands, so the request can reach somewhere else while carrying the credentials the scope was meant to protect.

Most of these servers keep logs, but the logs record the call as typed. A log full of receipts can still hide the fault if every receipt records the wrong side. His fix is for the receipt to carry the resolved target, the final address after substitution and clean-up, and he extends Evgenii's count: of the records that carry evidence, how many show what was declared and how many show what was executed. He offers a write-up of the pattern without vendor names.

Use cases I1 I3 I5 I10   Changes Row T10 now requires the resolved target in receipts.

Byte equality settles only what was declared

Kenne Ives; message

Kenne agrees that second sets of test vectors belong in the conformance folder, each checked first against the published examples in RFC 8785 so that no set depends on another implementation. Following Anas, the comparison subsection will say that identical bytes settle only the comparison of what was declared, and will name receipts bound to the executed effect as the open extension. Kenne will join calls when comparison semantics is on the agenda.

Use cases S6 I5   Changes No new row; wording for the §4 subsection.