Beyond Model Rankings: Evaluating AI Payment Agents with APort Vault
Uchi Uchibeke
What it found
- Rank association across replay tracks
- +0.673 (L2) to +0.967 (L5)
- Spearman over fourteen models, complete-case cohorts, session-bootstrap intervals
- Rates and rankings need not change together
- Level 2: 0.40% → 3.18%
- +2.77 percentage points, 95% interval [+1.68, +3.84], at a rank correlation of +0.673
- A stable total hides moving decisions
- Level 4: net 69, but 355 pairs disagree
- 212 request under one track only, 143 under the other
Abstract
We examine how adversarial replay changes model rankings and payment-request rates in APort Vault. This post hoc study uses fourteen models and complete-case cohorts from a frozen archive of 118,756 model-alone evaluations. At Level 2, on 690 prompts from 243 source sessions, payment requests occur in 39 of 9,660 final-user-turn evaluations and 307 of 9,660 capped source-user-turn evaluations: 0.40% against 3.18%. The corresponding Spearman rank correlation is +0.673, with a session-bootstrap 95% interval of [+0.48, +0.80]. At Level 4 the correlation is +0.913 while the aggregate count moves only from 7,193 to 7,124; of 9,030 matched pairs, 355 disagree while the net change is 69. Restricting to multi-user-turn sources reduces the correlations to +0.607 at Level 2 and +0.680 at Level 4. Positive rank association can coexist with substantial changes in action frequency.
Data and code
- Dataset on Hugging Face — every figure regenerates from the released rows
Related
- APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
4,371 human-authored attacks replayed against 14 models with and without a deterministic pre-action authorization check. 225,964 evaluations.
Cite
@misc{uchibeke2026rankings,
title = {Beyond Model Rankings: Evaluating {AI} Payment Agents with {APort} Vault},
author = {Uchibeke, Uchi},
year = {2026},
archivePrefix= {arXiv},
primaryClass = {cs.CR}
}