What the tests are#
Each case is one complete request to the check: a Grant, a Proof, the action the business submitted, and the exact state the check sees (the organization's DID document, the signed revocation list, which one-time values have already been used, whether a store is unavailable). With it comes the expected answer: the decision, the reason, the result of every check (24 in v0.2, up to 33 in v0.3), and the signed receipt the reference engine produces for that input.
The v0.2 cases cover an ALLOW for each action type and at least one DENY for every reason code in the v0.2 specification. That includes the ones that are easy to get subtly wrong: a replayed Proof, a revoked Grant, a revocation store that can't be reached, an unknown constraint, a Grant signed with alg: none, and a target URL that isn't in canonical form.
Three sets sit beside the cases. They cover the places where two correct-looking implementations most often disagree by one byte:
- htu: 32 checks of URL canonicalization (RFC 9449 §4.3). Uppercase schemes and hosts, default ports, queries and fragments, and inputs that must be rejected outright.
- jcs: 7 checks of RFC 8785 canonical JSON and the resulting
action_hash, including key ordering by UTF-16 code units and surrogate pairs. - jti: 12 checks of the one-time-value format.
The v0.3 set, run with --v3, adds cases for budgets, seller-side actions, dependencies between decisions and an organization's standing, plus a 49-check set for the sealed-commitment framing.
Everything is deterministic. The keys are fixed Ed25519 test keys, the clock is fixed at 2026-09-10T12:00:00.000Z, and every name is a non-production example (acme.example, test.delegus.example). Ed25519 signatures are deterministic, so an implementation that follows the specification reproduces each receipt byte for byte, not just "the same decision".
Run them locally#
npx @delegus/conformanceThat runs every case and every set against the reference engine and prints the result:
PASS: 40/40 vectors (40 receipts byte-identical), 3 sets passed, 0 skipped, 40/40 committed receipts re-verify offlineAnd the v0.3 set:
npx @delegus/conformance --v3PASS: 50/50 vectors (50 receipts byte-identical), 50/50 committed receipts checkThe last clause is a separate check. The receipt committed in each case is verified with its committed key and its protocol result is re-derived offline, independently of the implementation under test. The cases can't quietly drift from what they claim.
To test your own JavaScript implementation, point the runner at a module exporting the engine interface:
npx @delegus/conformance --impl ./my-verifier.js--only <id,id> runs a subset and --json prints the full report. The package is published on npm under the Apache-2.0 license.
Run them against an HTTP endpoint#
Most verifiers won't be written in JavaScript. For those, the runner speaks a small HTTP contract: one endpoint, JSON in and out. Every request is a POST carrying the wire version and the fingerprint of the test set, plus one operation:
{
"v": 1, 1
"vectors": "sha256:…", 2
"op": "canonicalize", 3
"inputs": ["{\"b\":1,\"a\":2}"] 4
}
- The wire version.
- A hash over every case and set file, so a report always says exactly which tests it ran.
- One of:
describe,evaluate,normalizeHtu,canonicalize,sha256,isJti. - For
canonicalize, the original JSON text, so your own parser handles the numbers.
Your endpoint answers {"results": [...]} for the batched set operations and an evaluation (decision, reason, both results, every check, and optionally the receipt) for evaluate. A full run is about forty requests.
The contract fails closed, deliberately:
- Every operation answers HTTP 200. An operation you haven't implemented answers
{"unsupported": true}, and only then is its set skipped. - A 404, any other error, a timeout, a body that isn't JSON, the wrong number of results, or a result of the wrong type is a failure, never a skip. A mistyped URL can't turn into a green report.
evaluateis never optional.- Answers are matched to questions by position, and
sha256is asked over your own canonical outputs, so each set checks your implementation end to end rather than ours.
Try it against your endpoint locally first:
npx @delegus/conformance --endpoint https://verifier.example/conformanceThe hosted runner#
The same run is available as a service, so you don't have to install anything. Start it with any Delegus API key:
POST https://api.delegus.ai/conformance/runs
{"endpoint": "https://verifier.example/conformance"}The answer is a 202 with a run_id. Read the result at GET https://api.delegus.ai/conformance/runs/{run_id}: the status is running, passed, failed or error, with the result of each case. The hosted runner runs the v0.2 set.
The runner calls your endpoint from our side, so it's fenced:
- public HTTPS endpoints only;
- no redirects;
- at most 50 requests and two minutes per run;
- one run at a time per account, and ten new runs per hour per API key.
It uses the same runner code as the local CLI, so a hosted result and a local result for the same endpoint and the same test set can't disagree.
What a report is, and isn't#
A report is Delegus's unsigned account of which cases your endpoint answered correctly, pinned to the fingerprint of the test set it ran. It is not a certificate, a badge or an endorsement.
It also covers exactly what the cases cover: the protocol layer. The four service checks (a registered organization, a verified domain, an uncompromised key, and a Proof not used before) depend on state the service holds. The cases let you show that your verifier evaluates them correctly given that state. They can't make your verifier the holder of that state. Passing every case means your verifier is a correct implementation of the published check. It doesn't make it Delegus.
That's the point. The check that decides whether an agent may act shouldn't be something you have to take on trust. Here are the tests; run them.
The conformance page · The specification, §5 and §15