Research

Every figure regenerates from a frozen snapshot, and the evaluation sets are released, so the numbers can be checked rather than taken.

preprint

Beyond Model Rankings: Evaluating AI Payment Agents with APort Vault

Changing how a recorded attack is replayed moves model rankings and action rates differently, and a stable total can hide decisions moving both ways.

Rank association across replay tracks
+0.673 (L2) to +0.967 (L5)
Rates and rankings need not change together
Level 2: 0.40% → 3.18%
A stable total hides moving decisions
Level 4: net 69, but 355 pairs disagree
publishedarXiv:2609.22076

APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport

4,371 human-authored attacks replayed against 14 models with and without a deterministic pre-action authorization check. 225,964 evaluations.

Unpermitted transfers, Levels 2 to 4
140 of 76,842 → 0 of 69,297
The zero is not deny-everything
25,370 payments executed behind the layer
Same inputs, fourteen models
809 of 1,293 Level 4 prompts elicited a request from every model