This is a focused narrative review of published research, not a new experiment or a systematic review. The sources below concern specific datasets and tasks. They do not measure SABHYA or current Instagram comment moderation.
The question
What does published work actually show about detecting harmful language in Hindi-English code-mixed and Romanized social-media text? The useful answer has two parts: researchers have built datasets and models for these distinct language conditions, and the evidence still leaves a large gap between a benchmark label and a dependable decision about a creator’s live comments.
Hinglish can switch language within a sentence. Romanized Hindi is written with Latin characters rather than Devanagari, often with variable spellings. A single comment may also include slang, emoji, sarcasm, quotation, or a reply whose meaning depends on another message. These are concrete inputs to review; they are not evidence that every English-language moderation system fails.
How we selected sources
On 22 September 2026, we reviewed primary papers in the ACL Anthology that directly describe Hindi-English code-mixing, Roman script, offensive or aggressive language labels, or language identification. We selected the five papers below because their published abstracts and descriptions state the data source and task clearly enough to compare scope. This is a small, purposive set, not a comprehensive search of all work through 2026. We did not combine results into a pooled accuracy estimate.
Our research methodology explains how we separate a published finding from an engineering inference and a SABHYA design statement. The source register records links and scope so a reader can inspect the originals.
Evidence at a glance
| Source | Data and task described by the authors | What it can support | Transfer limit |
|---|---|---|---|
| Bohra et al., 2018 | Hindi-English code-mixed tweets, annotated at word level for language and as hate or normal speech | Code-mixed hate-speech detection is a specific evaluated task | Twitter data and binary labels do not represent Instagram comments or SABHYA policy |
| Kumar et al., 2018 | An aggression-annotated Hindi-English code-mixed corpus | Aggression can be studied with a dedicated annotation scheme | Aggression labels are not interchangeable with a creator’s KEEP, REVIEW, HIDE decisions |
| Mathur et al., 2018 | Hindi-English code-switched tweets labelled non-offensive, abusive or hate speech | More than one harm category is possible | A tweet dataset does not establish outcomes for live Instagram reply threads |
| Nayak and Joshi, 2022 | Roman-script Hindi-English corpus and language-identification resources drawn from Twitter | Roman-script code-mixing warrants dedicated language processing | Language identification is not a harm verdict |
| Kumar et al., 2021 | A shared task for gender-biased and communal language in several Indian languages | Harm categories and languages differ across evaluation sets | Task labels and platforms do not give a production accuracy estimate for SABHYA |
What the studies establish
Code-mixing and script choice affect the task definition. Bohra et al. annotated word-level language alongside a hate-speech label. Nayak and Joshi built Roman-script Hindi-English and language-identification resources. Those design choices show why simply treating a Latin-script comment as ordinary English throws away relevant information. They do not show that language routing alone resolves whether a comment is abusive.
The target label changes the answer. One paper distinguishes hate speech from normal speech; another studies aggression; Mathur et al. separate non-offensive, abusive and hate-speech tweets. A creator may also care about spam, unwanted sexual attention, threats, or a remark that needs context. Scores on one label set cannot be silently reused for another policy.
The unit of context matters. These papers make bounded claims about their collected posts and annotations. An Instagram reply can refer to an earlier comment, a creator’s caption or a community-specific expression. We infer that a production workflow should preserve a path to inspect context when the input is ambiguous. The sources do not quantify how often that happens on Instagram.
What remains unmeasured here
No source in this review evaluates SABHYA on a current, representative sample of its intended Instagram use. We have not published a product benchmark, human-review agreement study, per-language false-positive rate, or measured improvement over Instagram’s native tools. Dataset years, collection methods, languages, scripts, platform norms and annotation policies differ. Those differences prevent a responsible claim that a reported research score transfers to today’s creator accounts.
A useful future evaluation would define the moderation policy before labelling, sample comments with appropriate permissions, record language/script and context, use more than one trained annotator for disputed cases, and report errors by route and harm category. It would separate model verdicts from whether a platform action was attempted and confirmed. The SABHYA system card explains the current product boundary; it is not evidence of completed testing.
What this means for product design
The following are SABHYA design statements, not research findings or measured performance claims. The product represents outcomes as KEEP, REVIEW and HIDE. REVIEW is a deliberate decision for uncertainty. Language routing informs processing but is not itself a toxicity finding. The private beta begins with human review while automatic moderation is validated. A failed provider response or unresolved case should not silently become a KEEP decision. See how SABHYA works, Hinglish moderation and the system card for the stated workflow and limitations.
Primary sources
- Bohra, A., Vijay, D., Singh, V., Akhtar, S. S. and Shrivastava, M. (2018). A Dataset of Hindi-English Code-Mixed Social Media Text for Hate Speech Detection. ACL Anthology, PEOPLES 2018.
- Kumar, R., Reganti, A. N., Bhatia, A. and Maheshwari, T. (2018). Aggression-annotated Corpus of Hindi-English Code-mixed Data. LREC 2018.
- Mathur, P., Shah, R., Sawhney, R. and Mahata, D. (2018). Detecting Offensive Tweets in Hindi-English Code-Switched Language. ACL Anthology, W-NUT 2018.
- Nayak, R. and Joshi, R. (2022). L3Cube-HingCorpus and HingBERT: A Code Mixed Hindi-English Dataset and BERT Language Models. WILDRE 2022.
- Kumar, R. et al. (2021). ComMA@ICON: Multilingual Gender Biased and Communal Language Identification Task at ICON-2021. ICON 2021.
Cite this research
Title: Hinglish and Romanized Indian Comment Moderation: What the Evidence Shows. Author/publisher: The Indian Alpha. Published: 22 September 2026. Type: Narrative evidence review, not a peer-reviewed paper or original benchmark. Canonical URL: https://theindianalpha.com/research/hinglish-romanized-indian-comment-moderation/
Suggested plain-text citation: The Indian Alpha (2026), “Hinglish and Romanized Indian Comment Moderation: What the Evidence Shows,” narrative evidence review, 22 September 2026, https://theindianalpha.com/research/hinglish-romanized-indian-comment-moderation/ (accessed [date]). Cite the underlying papers directly for their specific empirical findings.
Corrections: If a source is mischaracterized, contact The Indian Alpha through the contact page. This page will be updated when the evidence base or the product’s documented evaluation changes.