Banana Navy
Catalog FR
Lab · AI threat modeling · Detailed card

Risk scoring failure (false negatives, fragmented signals)

Letting a mis-judged attack through or blocking a legitimate call.

CardF14
Categoryrisk engine boundary (methodological blind spot)
Layers13 · Risk Scoring
SystemAI voicebot

Test the scoring against camouflaged attacks and keep a human on edge cases.

The threat

the risk engine under-rates an attack (false negative), over-rates a legitimate call (false positive) or merges fragmented attack signals poorly, letting a malicious action through.

Blind spotWhy classic frameworks miss it

scoring is often treated as a reliable box; its failure (merge flaw, thresholds) is a silent elevation vector nobody tests adversarially.

MitigationProposed approach

adversarial scoring tests (camouflaging an attack as benign signals), human-in-the-loop on edge cases, tracing of incoming signals.

The proposed control
no sensitive action passed on score alone without signal tracing.

Expected evidence
a fragmented campaign does not pass the risk engine without an alert.

SourcesReferences and public research

Public researchPublic research sources: MITRE ATLAS 2026.07 (verified technique mapping), OWASP GenAI (model abuse categories), and the public risk-voicebot (aivansoul/risk-voicebot) template defining the 20 checkpoints. No client registry data: generic card, no rating, no verdict.

Explore the 20 security layers

MITRE ATLAS 2026.07 · OWASP GenAI · risk-voicebot