Capability Exercise · capability.exercise

Somebody published a detection claim. This is how an outsider tests it and keeps the result.

Commit the challenge set and your expected answer for each one before you contact the provider. Then run it. The receipt records the detection rate we recomputed from your committed set, and it refuses to issue if you turn out to be affiliated with the provider you're testing.

Live in production

Why anyone needs this

Vendor detector accuracy is almost entirely self reported. Some vendors publish a detector that only recognizes their own marks. Some put it behind an email address. Some publish a number with no method attached, and some publish nothing at all and just say the words.

Meanwhile the published research keeps showing these schemes failing under ordinary processing. So when a buyer, a regulator or an expert witness asks "does the detector actually work," the honest answer today is that nobody outside the vendor has a defensible way to say.

The hard part isn't running the test. It's proving you didn't pick your challenge set after seeing which ones passed. That's what the precommitment is for.

How it works

01

Commit before you knock

You commit a digest over the challenge set and your declared expectation for each item. The receipt records that instant and refuses if it lands after the provider was first contacted.

02

Exercise the published interface

You invoke the capability at its published URL, with the invocation class recorded and the instant placed against a declared external time reference with a bounded drift.

03

We recompute the score

Per challenge outcomes, the detection rate in basis points and the conformance verdict are all recomputed by the service from your committed set. You cannot supply them.

Five real cases, against the production API

Three verify, including one where the capability could not be exercised at all because no interface was published. Two are refused.

The two refusals are the interesting part

A tuned run, where the challenge set was quietly changed after contact, is refused. So is an affiliated run, where the exerciser turns out not to be independent of the provider by a recomputed control set comparison. That second one matters: a receipt saying "we tested it and it passed" is worth nothing if we can't show who "we" was.

There is also a case for the situation that comes up most often in practice. The provider published a claim and no interface to test it with. That's recorded as an unexercisable outcome rather than a failure, because the honest finding is that the claim can't be checked from outside.

Every gate on a capability exercise (25 gates)
Gate
CAPABILITYEXERCISE_BOUNDARY_PRESENT
CAPABILITY_REFERENCE_DECLARED
CHALLENGE_PRECOMMITMENT_ORDER
CHALLENGE_SET_DIGEST_RECOMPUTE
CHALLENGE_SET_INTEGRITY
CONFORMANCE_VERDICT_RECOMPUTE
DETECTION_RATE_RECOMPUTE
DRIFT_BOUND_SATISFIED
EXERCISER_INDEPENDENCE_RECOMPUTE
EXERCISER_ROLE_DECLARED
EXERCISE_INSTANT_ORDERED
INVOCATION_CLASS_DECLARED
INVOCATION_OUTCOME_AGREEMENT
ISSUER_KEY_MATCH
NO_CHALLENGE_CONTENT_LEAK
OBSERVED_DRIFT_RECOMPUTE
OBSERVED_OUTPUT_BINDING
PAYLOAD_DIGEST_VALID
PER_CHALLENGE_OUTCOME_RECOMPUTE
PROVIDER_NONPARTICIPATION
PUBLISHED_CLAIM_BINDING
SALT_BINDING
SCHEMA
SIGNATURE_VALID
TIME_REFERENCE_INTEGRITY

What this receipt does not say

This is the honesty boundary carried inside every capability.exercise receipt, verbatim from the signed body.

This receipt attests only that a named exerciser, independent of the named provider by a recomputed control set comparison, committed a challenge set and a declared expectation for each challenge before the named capability was first contacted, then invoked that capability at the named URL through the stated invocation class at an instant placed against a declared external time reference with a declared drift bound, and that the per challenge outcomes, the detection rate, and the conformance verdict recorded here were recomputed by this service from the committed challenge set and the observed outputs rather than supplied by the caller. It does not attest that the capability is accurate, effective, or fit for any purpose. It does not attest that any marking scheme is robust. It does not attest that the provider complies with any law, regulation, or standard. It does not attest that the capability was available at any instant other than the one recorded. It does not attest that a different exerciser, network path, or challenge set would produce the same outputs. It does not attest that the provider's published claim is true or false beyond the declared expectation tested here. It does not attest that no equivalent capability exists anywhere else. It does not decide whether any content was generated by any system.

Endpoints

RouteWhat it does
POST /v1/verify/capability-exerciseVerify an exercise receipt against the precommitted challenge set
GET /v1/demo/capability-exerciseFive worked cases, two of them refused

Where this sits

Capability exercise measures somebody else's claim from outside. Capture commitment handles the other half of the audio question, whether a recording was cut. Put either receipt in a transparency checkpoint and the result is also pinned in time by an authority that has nothing to do with Hive.

Patent pending. Hive Civilization, The Hivery, Inc.