This document defines the conformance and interoperability framework for the Agent Trust Protocol (ATP) core specifications. It describes the per-item conformance suites, the structure of the vendor-neutral test vectors, and the audit-store integrity contract that an ATP audit/trust-registry implementation MUST satisfy. Its purpose is to let independent implementations demonstrate interoperability against the same deterministic vectors.

This document is an early draft work product of the Agent Trust Protocol (ATP) Community Group. It is not a W3C Standard nor on the W3C Recommendation track. It is published to seek review and contribution from Community Group members.

The reference implementation ships a strong internal conformance harness. The work this document tracks is publishing vendor-neutral, spec-bound vectors any implementation can run. Where a vector below is drawn from the reference harness it is marked as such; CG members are invited to contribute independent vectors.

As well as sections marked as non-normative, all authoring guidelines, diagrams, examples, and notes in this specification are non-normative. Everything else is normative.

The key words MUST, MUST NOT, SHOULD, and MAY are to be interpreted as described in BCP 14 [[RFC2119]] [[RFC8174]] when, and only when, they appear in all capitals.

Conformance Items

The ATP core surface is partitioned into the following conformance items. A conforming implementation MUST pass the suite for each item it claims to implement.

#ItemDefining specification
1did:atp — create / parse / resolvedid:atp DID Method
2Hybrid signatures — Ed25519 + ML-DSA-65did:atp DID Method
3Policy assertion — allow/deny, deny-by-defaultATP Trust
4Audit store / trust registry — integrityThis document, §Audit Integrity
5Privacy — selective disclosure (Merkle membership)ATP Privacy
6Trust scoring — anti-gaming (decay, confidence, endorser independence)ATP Trust

The reference implementation runs all items from a single entry point and exits non-zero on any failure, so it doubles as a release gate. Each item is a table-driven suite into which external vectors drop in as new rows.

Test-Vector Format

Vectors are published as JSON keyed by conformance item. Each vector is a deterministic known-answer test: a fixed input and the exact expected output (an identifier, a hash, a decision, or a boolean), computed independently of any one implementation. A conforming implementation MUST reproduce the expected output for every published vector for the items it claims.

The packaging of the vendor-neutral vector file (single bundle vs. per-item files, and the exact JSON schema) is open for CG discussion.

Anti-Gaming Vectors (Item 6)

These vectors exercise the farming and Sybil cases against the Trust-Scoring Function. They assert trust bands, per-capability confidence, and invariants — not point scores. No scalar scoring formula is asserted; the numeric mapping remains out of scope, so an implementation may satisfy these vectors under any formula consistent with the stated constraints.

Fixture A reuses the seeded example from the Trust specification verbatim and asserts that it MUST NOT reach a privileged or trusted band. Fixture B shows fewer but diverse, recent, receipt-backed interactions earning confidence bounded to the capabilities their evidence covers.

Both vectors are evaluated under the default profile parameters — a 90-day recency half-life and a minimum of three effective independent endorsers. An implementation using different values MUST state them when making a conformance claim, and the expected outputs below apply to the default profile.

VectorScenarioExpected
atp-antigaming-A-volume-farming 500 same-type low-risk interactions, one counterparty; 20 endorsements from one funding source with mutual edges Global band < TRUSTED; high confidence for the single low-risk capability only
atp-antigaming-B-diverse-verified 40 successful / 3 failed across 5 action types and 30 independent counterparties, receipt-backed; 6 independent endorsers Confidence scoped to the capabilities receipts cover; uncovered capabilities remain none
{
  "id": "atp-antigaming-A-volume-farming",
  "conformance_item": "scoring/anti-gaming/volume-farming",
  "input": {
    "identityVerified": true,
    "credentials": ["a", "b", "c", "d"],
    "successful": 500, "failed": 0,
    "endorsements": 20, "violations": 0,
    "interactions": {
      "distinct_action_types": 1,
      "action_type": "read_public_score",
      "action_risk_class": "low",
      "value_or_authority_change": false,
      "counterparty_independence": { "distinct_counterparties": 1, "cluster": "single" }
    },
    "endorser_independence": {
      "distinct_funding_sources": 1,
      "mutual_edges": true,
      "effective_independent_endorsers": 1
    },
    "recency": { "within_decay_window": true }
  },
  "expected": {
    "global_band_constraint": "< TRUSTED",
    "per_capability_confidence": [
      { "capability": "read_public_score", "confidence": "high", "authority_frame_required": true }
    ],
    "invariant_checks": [
      "global band is flat on raw interaction count",
      "endorsement contribution requires >= 3 effective independent endorsers",
      "no global-band elevation without capability-scoped, independent evidence"
    ]
  }
}
      
{
  "id": "atp-antigaming-B-diverse-verified",
  "conformance_item": "scoring/anti-gaming/evidence-quality",
  "input": {
    "identityVerified": true,
    "credentials": ["a", "b"],
    "successful": 40, "failed": 3,
    "endorsements": 6, "violations": 0,
    "interactions": {
      "distinct_action_types": 5,
      "failure_weighting": "by_severity",
      "counterparty_independence": { "distinct_counterparties": 30, "cluster": "diverse" },
      "outcomes_with_receipts": true,
      "receipt_shape": {
        "action_type": true, "authority_boundary": true, "counterparty": true,
        "timestamp": true, "outcome_hash": true, "recomputable": true
      },
      "capabilities_with_receipt_coverage": [
        "submit_signed_receipt", "verify_credential", "resolve_did"
      ]
    },
    "endorser_independence": {
      "distinct_funding_sources": 6,
      "mutual_edges": false,
      "effective_independent_endorsers": 6
    },
    "recency": { "within_decay_window": true }
  },
  "expected": {
    "per_capability_confidence": [
      { "capability": "submit_signed_receipt", "confidence": "high" },
      { "capability": "verify_credential",     "confidence": "high" },
      { "capability": "resolve_did",           "confidence": "medium" },
      { "capability": "read_public_score",     "confidence": "none" },
      { "capability": "post_endorsement",      "confidence": "none" }
    ],
    "invariant_checks": [
      "band moves on independent, recent, recomputable evidence",
      "failures are included and weighted by severity (do not raise the band)",
      "confidence is scoped to capabilities with receipt coverage"
    ]
  }
}
      

Fixture B's recomputable receipts resolve to audit events whose hashes a verifier recomputes, as defined in §Audit-Store Integrity. Whether the scoring layer MUST require that outcomes resolve to conformance-layer audit events, or whether that is a separate profile, is one of the open questions in the capability-scoped confidence discussion (issue #4).

These vectors express confidence per capability. The broader question of whether ATP replaces the single global band with capability-scoped confidence as the authorization input is a distinct architectural decision under discussion in issue #4, and is not settled by adopting these vectors. Until that decision is recorded, per-capability confidence reported here sits alongside the global band rather than replacing it.

Audit-Store Integrity

An ATP audit store is an append-only hash chain. A conforming implementation MUST guarantee two properties:

  1. an intact chain verifies as valid; and
  2. any tamper — mutating an entry's content, reordering entries, or breaking the previousHash linkage — is detected (verification returns false).

Integrity verification MUST recompute each event's SHA-256 hash over the canonical, storage-independent field set and check previousHash continuity from a genesis value of 64 zero hex characters. The hash MUST be computed over exactly these fields, in canonical (recursively key-sorted) JSON:

{
  "id":           "...",
  "timestamp":    "...",
  "source":       "...",
  "action":       "...",
  "resource":     "...",
  "actor":        "...",
  "details":      { /* ... */ },
  "previousHash": "<hex sha256 of prior event, or 64×'0' for genesis>"
}
      

Storage-managed metadata such as a nonce or a backend-assigned block number MUST NOT be part of the integrity hash: backends do not round-trip these fields uniformly, so including them would make recomputed hashes diverge and falsely report tampering. Event ordering is bound cryptographically by previousHash.

Integrity verification MUST perform its own content verification and MUST NOT trust a storage-supplied "chain valid" verdict, which typically checks linkage only and would miss content tampering.

Conformance Claims

An implementation claiming ATP conformance MUST state which items (1–6) it implements and MUST pass every published vector for those items. A claim of "ATP Core" conformance requires items 1, 2, and 4.

Item 6 (anti-gaming) is not part of the "ATP Core" set in this version. Its requirements are newly adopted and its parameters are still provisional; requiring it for a core claim would invalidate existing implementations before the profile has settled. An implementation claiming item 6 MUST state the profile parameters it uses (Anti-Gaming Vectors). Whether item 6 joins the core set in a future version is left open.

Whether to define named conformance profiles (e.g. "Identity-only" vs. "Full ATP") is open for CG discussion.