This is a proposed vocabulary for research planning. It is not a report of observed SABHYA error rates.
| Error family | Working definition | Example evaluation question |
|---|---|---|
| False positive | A comment is flagged for intervention even though the applicable policy would allow it. | Did quotation or benign slang trigger the wrong outcome? |
| False negative | A comment requiring intervention is left without the intended handling. | Did spelling variation or obfuscation change detection? |
| Context error | The decision ignores relevant conversational, target, or creator-policy context. | Would adjacent replies change interpretation? |
| Language-route error | A comment is assigned to an unsuitable language-processing route. | Does script or code-mixing change route selection? |
| Uncertainty error | The system fails to surface a case that should be reviewed. | Was an ambiguous case silently forced into a binary outcome? |
| Action-state error | The displayed operational state does not match the platform response. | Is a decision incorrectly represented as a completed platform action? |
Any future benchmark should define labels, sampling, annotation adjudication, privacy controls, subgroup reporting, and limitations before publishing quantitative findings.