Pisama
Benchmarks/Detector scoreboard

Detector Calibration Scoreboard

Every Pisama detector, calibrated on golden datasets and reported per-failure-mode. Hyperscaler agent platforms advertise “anomaly detection” without publishing per-detector F1. We publish the whole table.

Detectors measured
52
of 87 in the registry
Mean F1
0.753
range 0.0001.000
Production grade
6
7 beta · 18experimental · 21 failing
Last calibrated
2026-06-14
external-only lane (real traces)

Externally validated at production grade: real-trace F1 0.80 or higher, precision 0.70 or higher, 30 or more external traces, external-grounded thresholds, and no per-difficulty blind spot (capability registry, external-only lane, 2026-06-14).

Tier
persona_driftexperimental1.000100.0%100.0%13
workflowfailing1.000100.0%100.0%6
retrieval_qualityfailing1.000100.0%100.0%5
n8n_complexityfailing1.000100.0%100.0%2
openclaw_tool_abusefailing1.000100.0%100.0%4
openclaw_spawn_chainfailing1.000100.0%100.0%4
openclaw_channel_mismatchfailing1.000100.0%100.0%4
openclaw_sandbox_escapefailing1.000100.0%100.0%4
convergencebeta1.000100.0%100.0%16
citationfailing1.000100.0%100.0%2
openclaw_session_loopfailing1.000100.0%100.0%4
consensus_collapseexperimental1.000100.0%100.0%10
chunk_relevancefailing1.000100.0%100.0%4
chunk_attributionfailing1.000100.0%100.0%4
deceptionexperimental1.000100.0%100.0%13
impersonation_riskexperimental1.000100.0%100.0%13
scope_escalationexperimental1.000100.0%100.0%13
reward_hackingexperimental1.000100.0%100.0%8
synthesis_failureproduction0.96597.9%95.0%200
corruptionexperimental0.941100.0%88.9%14
rag_poisoningbeta0.94188.9%100.0%16
decompositionproduction0.94098.4%90.0%104
silent_cascadeproduction0.93798.9%89.0%200
specificationbeta0.92692.6%92.6%113
injectionbeta0.87577.8%100.0%15
role_usurpation_execexperimental0.857100.0%75.0%13
sycophancyexperimental0.857100.0%75.0%8
over_refusalbeta0.83591.5%76.8%112
completionproduction0.82894.2%73.9%106
contextproduction0.81293.2%71.9%62
hallucinationproduction0.80774.2%88.5%113
openclaw_elevated_riskfailing0.80066.7%100.0%4
groundingbeta0.75262.1%95.3%78
role_usurpationbeta0.692100.0%52.9%26
n8n_errorfailing0.66750.0%100.0%3
delegationfailing0.66766.7%66.7%7
context_precisionfailing0.66750.0%100.0%4
multi_agent_contagionexperimental0.66775.0%60.0%10
loopexperimental0.60079.0%48.4%49
under_refusalexperimental0.57892.3%42.1%113
redundant_delegation_conflictexperimental0.57395.3%41.0%200
communicationexperimental0.56875.0%45.6%61
jailbreak_complianceexperimental0.50795.0%34.5%112
output_validationexperimental0.500100.0%33.3%10
coordinationexperimental0.44445.2%43.8%113
routingexperimental0.44450.0%40.0%10
derailmentfailing0.26785.7%15.8%72
role_usurpation_canonicalfailing0.267100.0%15.4%22
withholdingfailing0.0000.0%0.0%2
n8n_resourcefailing0.0000.0%0.0%2
specification_compliancefailing0.0000.0%0.0%4
analytical_semanticsfailing0.0000.0%0.0%5

Why publish this

Generic anomaly detection is becoming a feature in every hyperscaler agent platform. That makes “we have anomaly detection” a check-the-box claim, not a differentiator.

Pisama’s position is the opposite: a structured failure taxonomy with calibrated detectors, each tuned and reported on a per-mode basis. Coordination failure does not look like grounding failure does not look like persona drift, and the eval surface should reflect that.

Numbers above come from the capability registry (external-only lane): real traces from public benchmarks and production integrations, cross-validated per detector. Tier is the registry readiness; the production gate is spelled out under the summary cards. Failing detectors stay in the table, since publishing only the winners would misrepresent the measured population.

Source code: github.com/Pisama-AI/pisama. Snapshot generated by backend/scripts/generate_scoreboard_snapshot.py.

© 2026 Pisama. All rights reserved.