Detector Calibration Scoreboard
Every Pisama detector, calibrated on golden datasets and reported per-failure-mode. Hyperscaler agent platforms advertise “anomaly detection” without publishing per-detector F1. We publish the whole table.
Externally validated at production grade: real-trace F1 0.80 or higher, precision 0.70 or higher, 30 or more external traces, external-grounded thresholds, and no per-difficulty blind spot (capability registry, external-only lane, 2026-06-14).
| Tier | |||||
|---|---|---|---|---|---|
| persona_drift | experimental | 1.000 | 100.0% | 100.0% | 13 |
| workflow | failing | 1.000 | 100.0% | 100.0% | 6 |
| retrieval_quality | failing | 1.000 | 100.0% | 100.0% | 5 |
| n8n_complexity | failing | 1.000 | 100.0% | 100.0% | 2 |
| openclaw_tool_abuse | failing | 1.000 | 100.0% | 100.0% | 4 |
| openclaw_spawn_chain | failing | 1.000 | 100.0% | 100.0% | 4 |
| openclaw_channel_mismatch | failing | 1.000 | 100.0% | 100.0% | 4 |
| openclaw_sandbox_escape | failing | 1.000 | 100.0% | 100.0% | 4 |
| convergence | beta | 1.000 | 100.0% | 100.0% | 16 |
| citation | failing | 1.000 | 100.0% | 100.0% | 2 |
| openclaw_session_loop | failing | 1.000 | 100.0% | 100.0% | 4 |
| consensus_collapse | experimental | 1.000 | 100.0% | 100.0% | 10 |
| chunk_relevance | failing | 1.000 | 100.0% | 100.0% | 4 |
| chunk_attribution | failing | 1.000 | 100.0% | 100.0% | 4 |
| deception | experimental | 1.000 | 100.0% | 100.0% | 13 |
| impersonation_risk | experimental | 1.000 | 100.0% | 100.0% | 13 |
| scope_escalation | experimental | 1.000 | 100.0% | 100.0% | 13 |
| reward_hacking | experimental | 1.000 | 100.0% | 100.0% | 8 |
| synthesis_failure | production | 0.965 | 97.9% | 95.0% | 200 |
| corruption | experimental | 0.941 | 100.0% | 88.9% | 14 |
| rag_poisoning | beta | 0.941 | 88.9% | 100.0% | 16 |
| decomposition | production | 0.940 | 98.4% | 90.0% | 104 |
| silent_cascade | production | 0.937 | 98.9% | 89.0% | 200 |
| specification | beta | 0.926 | 92.6% | 92.6% | 113 |
| injection | beta | 0.875 | 77.8% | 100.0% | 15 |
| role_usurpation_exec | experimental | 0.857 | 100.0% | 75.0% | 13 |
| sycophancy | experimental | 0.857 | 100.0% | 75.0% | 8 |
| over_refusal | beta | 0.835 | 91.5% | 76.8% | 112 |
| completion | production | 0.828 | 94.2% | 73.9% | 106 |
| context | production | 0.812 | 93.2% | 71.9% | 62 |
| hallucination | production | 0.807 | 74.2% | 88.5% | 113 |
| openclaw_elevated_risk | failing | 0.800 | 66.7% | 100.0% | 4 |
| grounding | beta | 0.752 | 62.1% | 95.3% | 78 |
| role_usurpation | beta | 0.692 | 100.0% | 52.9% | 26 |
| n8n_error | failing | 0.667 | 50.0% | 100.0% | 3 |
| delegation | failing | 0.667 | 66.7% | 66.7% | 7 |
| context_precision | failing | 0.667 | 50.0% | 100.0% | 4 |
| multi_agent_contagion | experimental | 0.667 | 75.0% | 60.0% | 10 |
| loop | experimental | 0.600 | 79.0% | 48.4% | 49 |
| under_refusal | experimental | 0.578 | 92.3% | 42.1% | 113 |
| redundant_delegation_conflict | experimental | 0.573 | 95.3% | 41.0% | 200 |
| communication | experimental | 0.568 | 75.0% | 45.6% | 61 |
| jailbreak_compliance | experimental | 0.507 | 95.0% | 34.5% | 112 |
| output_validation | experimental | 0.500 | 100.0% | 33.3% | 10 |
| coordination | experimental | 0.444 | 45.2% | 43.8% | 113 |
| routing | experimental | 0.444 | 50.0% | 40.0% | 10 |
| derailment | failing | 0.267 | 85.7% | 15.8% | 72 |
| role_usurpation_canonical | failing | 0.267 | 100.0% | 15.4% | 22 |
| withholding | failing | 0.000 | 0.0% | 0.0% | 2 |
| n8n_resource | failing | 0.000 | 0.0% | 0.0% | 2 |
| specification_compliance | failing | 0.000 | 0.0% | 0.0% | 4 |
| analytical_semantics | failing | 0.000 | 0.0% | 0.0% | 5 |
Why publish this
Generic anomaly detection is becoming a feature in every hyperscaler agent platform. That makes “we have anomaly detection” a check-the-box claim, not a differentiator.
Pisama’s position is the opposite: a structured failure taxonomy with calibrated detectors, each tuned and reported on a per-mode basis. Coordination failure does not look like grounding failure does not look like persona drift, and the eval surface should reflect that.
Numbers above come from the capability registry (external-only lane): real traces from public benchmarks and production integrations, cross-validated per detector. Tier is the registry readiness; the production gate is spelled out under the summary cards. Failing detectors stay in the table, since publishing only the winners would misrepresent the measured population.
Source code: github.com/Pisama-AI/pisama. Snapshot generated by backend/scripts/generate_scoreboard_snapshot.py.