Hinglish comments may mix Hindi and English in one sentence. Romanized Hindi uses Latin letters, often with inconsistent spelling. A comment can also depend on the post, a previous reply, a joke, or the creator’s stated boundaries. Those conditions make a one-word rule a poor substitute for review. Our evidence review describes what published datasets actually support and where transfer to Instagram remains unmeasured.

Language route, moderation verdict and platform action

These are three separate questions. A system might recognize a comment as likely Hindi-English code-mixed text. It still has to assess whether the comment violates a policy. Even if a reviewer chooses HIDE, the workflow must separately confirm what Instagram did. A missing signal at any stage should be visible rather than silently converted to KEEP.

Three separate stages: likely language route, policy verdict, and confirmed platform action; uncertainty can lead to human review.

The moderation error taxonomy describes how mistakes can arise at each boundary.

Synthetic examples that need different treatment

These are invented examples for explanation. They are not labelled test data and do not demonstrate SABHYA performance.

Comment fragmentWhy a keyword or language tag is insufficientSensible first step
“Yeh idea kaafi useful tha”Romanized Hindi can express ordinary praise.KEEP if the surrounding context agrees.
“That word was used in the abuse I received”A harmful word may be quoted rather than directed at someone.REVIEW the surrounding thread.
“Wah, kya great advice…”Literal praise and sarcasm can share words.REVIEW when the context matters.
“Tu toh genius hai 😂”The same phrase and emoji could be friendly teasing or a targeted put-down.REVIEW the relationship and reply context before hiding.
“Tum log aise hi ho”It is unclear who “tum log” refers to without the thread; the target could change the policy decision.REVIEW the parent comment and caption.
“Send your address, I know where you live”The risk is the directed message, not its language route.Escalate under the creator’s threat policy; consider reporting.

Informal spelling adds another difficulty: “bahut”, “bohot” and “bahot” may be variants of the same ordinary word. A fixed keyword list can miss variants, while a broad pattern may catch harmless words. This illustrates input variation, not a measured error rate. The published sources discuss code-mixed and Romanized data, but none measures these examples in SABHYA.

The same fragment can also change meaning when its target changes. “Great job” under a tutorial may be praise; repeated as a reply to someone describing harassment, with mocking context, it may be hostile. No spelling normalization or language label resolves that difference on its own. A reviewer should see the relevant reply and creator policy, then leave the reason for the decision visible to the account owner.

A review process for uncertain comments

  1. Preserve the original text and only the minimum context a reviewer needs.
  2. Record a likely language/script route as a clue, not as a toxicity conclusion.
  3. Check whether the text is directed at someone, quoted, joking, supportive, or part of a reply chain.
  4. Apply the creator’s written policy. Use REVIEW when the available evidence cannot justify KEEP or HIDE.
  5. Record the moderation verdict and, separately, whether any Instagram visibility action succeeded.
  6. Revisit false positives and false negatives with privacy protected. For a credible threat or exposed personal information, use a human escalation path and Instagram reporting.

What SABHYA says it does

SABHYA is a private-beta Instagram comment moderation product built by The Indian Alpha. Its stated outcomes are KEEP, REVIEW and HIDE. REVIEW is intended to be a real destination for uncertainty. Its design separates a language route from a policy verdict and from the platform action state. We have not published a product benchmark, a per-language error rate or a measured improvement over Instagram native controls. Read the system card and human-control approach for current limits.

Start with the research evidence and primary sources. For an account-level workflow, see the Instagram moderation guide and native-controls comparison.