← Research
preprintcs.CR, cs.AI

Beyond Model Rankings: Evaluating AI Payment Agents with APort Vault

Uchi Uchibeke

What it found

Rank association across replay tracks
+0.673 (L2) to +0.967 (L5)
Spearman over fourteen models, complete-case cohorts, session-bootstrap intervals
Rates and rankings need not change together
Level 2: 0.40% → 3.18%
+2.77 percentage points, 95% interval [+1.68, +3.84], at a rank correlation of +0.673
A stable total hides moving decisions
Level 4: net 69, but 355 pairs disagree
212 request under one track only, 143 under the other

Abstract

We examine how adversarial replay changes model rankings and payment-request rates in APort Vault. This post hoc study uses fourteen models and complete-case cohorts from a frozen archive of 118,756 model-alone evaluations. At Level 2, on 690 prompts from 243 source sessions, payment requests occur in 39 of 9,660 final-user-turn evaluations and 307 of 9,660 capped source-user-turn evaluations: 0.40% against 3.18%. The corresponding Spearman rank correlation is +0.673, with a session-bootstrap 95% interval of [+0.48, +0.80]. At Level 4 the correlation is +0.913 while the aggregate count moves only from 7,193 to 7,124; of 9,030 matched pairs, 355 disagree while the net change is 69. Restricting to multi-user-turn sources reduces the correlations to +0.607 at Level 2 and +0.680 at Level 4. Positive rank association can coexist with substantial changes in action frequency.

Data and code

Related

Cite

@misc{uchibeke2026rankings,
  title        = {Beyond Model Rankings: Evaluating {AI} Payment Agents with {APort} Vault},
  author       = {Uchibeke, Uchi},
  year         = {2026},
  archivePrefix= {arXiv},
  primaryClass = {cs.CR}
}