Pisama versus the frontier, on independent benchmarks.
Two third-party academic benchmarks measuring the same question from different angles: did a failure happen? (TRAIL) and which agent failed, at which step?(Who&When, ICML 2025). Same traces, same labels.
Joint accuracy (heuristic detectors only, Tier 1 to 3, no LLM judge): detector predictions matching ground-truth labels on the full TRAIL set (148 traces, 841 failures).
148 traces · 841 labelled failures · frontier numbers from TRAIL paper
Best frontier shown is GPT-5.5 (current). The earlier GPT-5.4 scored marginally higher at 11.9%.
Counts are the capability registry, the same source as the per-detector scoreboard: 87 detectors total, 52 measured on the real-trace (external-only) lane, 6 externally validated at production grade. The mean F1 above is over that production-grade set; the full measured set spans failing to production. Externally validated at production grade: real-trace F1 0.80 or higher, precision 0.70 or higher, 30 or more external traces, external-grounded thresholds, and no per-difficulty blind spot (capability registry, external-only lane, 2026-06-14). Per-detector F1, precision, recall, and confidence intervals on the full scoreboard. MAST-mapped view on the taxonomy page.