Roblox Opens Three AI Safety Models for Child-Protection Teams
The models give other services tools to test and adapt, but Roblox’s own metrics and research on messages that slipped through moderation show automation remains an incomplete safeguard.
Listen to this story
The audio brief
Story brief
3 key pointsRoblox is releasing three updated safety models and an evaluation dataset through the ROOST Model Community, extending child-protection tooling to other online services. The package covers PII and off-platform contact detection, child-endangerment prioritization, and real-time voice moderation. Results are promising but bounded: Roblox reports a PII F1 increase to 90.52, while voice recall is 61% at a 1%...
- 01
PII Classifier Version 2 expands language support from 17 to 189 and evaluates message sequences, including misspellings, fragmented details, and coded language.
- 02
Sentinel Version 2 adds configurable risk scoring and explanations to help reviewers prioritize possible child-endangerment cases.
- 03
Voice Safety Classifier Version 3 covers 30 languages and eight violation categories, but its reported 61% recall means violations can be missed.
Roblox is contributing three updated AI safety models to the ROOST Model Community, making tools for detecting personal-information requests, possible child endangerment and unsafe voice chat available beyond its own platform. The release gives online services a starting point, not a finished safety operation.
The August 19 release includes PII Classifier Version 2, Sentinel Version 2 and Voice Safety Classifier Version 3. Roblox is also releasing an evaluation dataset for platforms to test their own safety systems. Roblox joined ROOST as a founding member in 2025 alongside Google, OpenAI, Discord and others.
Reading a conversation, not just a phrase
PII Classifier Version 2 examines requests for personally identifiable information and efforts to move users to another platform. Roblox says it evaluates messages together, rather than in isolation, to account for misspellings, fragmented contact details and coded language. It says the update expanded language support from 17 to 189 languages.
- Sentinel Version 2 adds options for scoring suspicious behavior and provides more information about why an interaction received a score.
- Roblox says Sentinel looks for early signs of possible child endangerment so human reviewers can prioritize interactions that may require action.
- Voice Safety Classifier Version 3 analyzes speech for policy violations in real time and covers 30 languages and eight violation categories.
Signals still need people and policies
Roblox says Sentinel’s early detection accounted for nearly 70% of the child-endangerment cases it identified in the 12 months ending August 7, 2026. That is an operational result from Roblox, rather than evidence that the model will perform the same way on another service.
Voice moderation shows a separate constraint. Roblox reports 61% recall for Version 3 across its supported languages at a 1% false-positive rate. Under that setting, the system can miss violations. The company says its earlier open-source voice classifier has been downloaded more than 72,000 times since 2024, but a download does not establish deployment, tuning or reviewer capacity.
A gap the release does not erase
Independent researchers analyzed more than 2 million Roblox chat messages and found examples of grooming, sexualization of minors, bullying, violence and sensitive-information sharing that escaped existing moderation. The new models target some of those weaknesses, particularly patterns that emerge across messages. Whether they improve protection elsewhere depends on platforms adopting them, adapting them to their services and acting on the alerts.
Sources
- foxnews.comRoblox shares AI child-safety tools parents should know
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.