Roblox Opens Three AI Safety Models for Child-Protection Teams

The models give other services tools to test and adapt, but Roblox’s own metrics and research on messages that slipped through moderation show automation remains an incomplete safeguard.

By 2 min read
Roblox Opens Three AI Safety Models for Child-Protection Teams
Roblox Opens Three AI Safety Models for Child-Protection Teams

Listen to this story

The audio brief

About 1:36
0:001:36
Read transcript
Roblox is opening three of its child-safety AI models to other online services, along with an evaluation dataset for testing them. The release runs through the ROOST Model Community and covers three different problems: detecting requests for personal information or attempts to move a conversation off-platform, prioritizing possible child-endangerment cases, and moderating voice chat in real time. PII Classifier Version 2 looks at message sequences instead of isolated phrases. That lets it account for misspellings, fragmented contact details, and coded language. Roblox says it expanded language support from 17 to 189 languages, and reports that the model’s F1 score rose from 63.41 to 90.52. Sentinel Version 2 adds configurable risk scoring and explanations, so human reviewers can focus first on interactions that may need action. Roblox says Sentinel detected nearly 70 percent of the child-endangerment cases it identified during the 12 months ending August 7, 2026. That is Roblox’s own operational result, not a guarantee for another platform. The voice model supports 30 languages and eight violation categories, but its reported recall is 61 percent at a one-percent false-positive rate. In plain terms, it can still miss violations. That limitation matters because independent researchers found grooming, sexualization of minors, violence, bullying, and sensitive-information sharing in more than two million Roblox chat messages that escaped existing moderation. The open question is whether other services will tune these models locally, staff human review, and act consistently on the alerts.

Story brief

3 key points

Roblox is releasing three updated safety models and an evaluation dataset through the ROOST Model Community, extending child-protection tooling to other online services. The package covers PII and off-platform contact detection, child-endangerment prioritization, and real-time voice moderation. Results are promising but bounded: Roblox reports a PII F1 increase to 90.52, while voice recall is 61% at a 1%...

  1. 01

    PII Classifier Version 2 expands language support from 17 to 189 and evaluates message sequences, including misspellings, fragmented details, and coded language.

  2. 02

    Sentinel Version 2 adds configurable risk scoring and explanations to help reviewers prioritize possible child-endangerment cases.

  3. 03

    Voice Safety Classifier Version 3 covers 30 languages and eight violation categories, but its reported 61% recall means violations can be missed.

Roblox is contributing three updated AI safety models to the ROOST Model Community, making tools for detecting personal-information requests, possible child endangerment and unsafe voice chat available beyond its own platform. The release gives online services a starting point, not a finished safety operation.

The August 19 release includes PII Classifier Version 2, Sentinel Version 2 and Voice Safety Classifier Version 3. Roblox is also releasing an evaluation dataset for platforms to test their own safety systems. Roblox joined ROOST as a founding member in 2025 alongside Google, OpenAI, Discord and others.

Reading a conversation, not just a phrase

PII Classifier Version 2 examines requests for personally identifiable information and efforts to move users to another platform. Roblox says it evaluates messages together, rather than in isolation, to account for misspellings, fragmented contact details and coded language. It says the update expanded language support from 17 to 189 languages.

  • Sentinel Version 2 adds options for scoring suspicious behavior and provides more information about why an interaction received a score.
  • Roblox says Sentinel looks for early signs of possible child endangerment so human reviewers can prioritize interactions that may require action.
  • Voice Safety Classifier Version 3 analyzes speech for policy violations in real time and covers 30 languages and eight violation categories.

Signals still need people and policies

Roblox says Sentinel’s early detection accounted for nearly 70% of the child-endangerment cases it identified in the 12 months ending August 7, 2026. That is an operational result from Roblox, rather than evidence that the model will perform the same way on another service.

Voice moderation shows a separate constraint. Roblox reports 61% recall for Version 3 across its supported languages at a 1% false-positive rate. Under that setting, the system can miss violations. The company says its earlier open-source voice classifier has been downloaded more than 72,000 times since 2024, but a download does not establish deployment, tuning or reviewer capacity.

A gap the release does not erase

Independent researchers analyzed more than 2 million Roblox chat messages and found examples of grooming, sexualization of minors, bullying, violence and sensitive-information sharing that escaped existing moderation. The new models target some of those weaknesses, particularly patterns that emerge across messages. Whether they improve protection elsewhere depends on platforms adopting them, adapting them to their services and acting on the alerts.

Sources

  1. foxnews.comRoblox shares AI child-safety tools parents should know

Loading discussion...