Scale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation Gap
The study found lower harmful-response scores in Korean prompts, but also more refusals of harmless requests—complicating claims that a model is simply safer in another language.
Listen to this story
The audio brief
Story brief
3 key pointsScale AI and the Korea AI Safety Institute introduced ROK-FORTRESS, a benchmark designed to distinguish language effects from local geopolitical context in AI safety testing. Across 14 models and 1,235 tasks, Korean-language prompts changed harmful-response rates by about 10 percentage points, versus roughly four points for switching U.S. references to Korean ones. The result is not a straightforward safety win: 12...
- 01
The benchmark independently varies English/Korean wording and U.S./Korean entities, institutions, and operational details.
- 02
Removing role-play and emotional wrappers mostly eliminated the Korean advantage, suggesting framing can influence apparent safety differences.
- 03
Five open-source frontier models were more likely to comply in Korean; proprietary models remained modestly safer.
A model that appears safer in Korean may also be less willing to answer harmless Korean questions. That is the central tension in ROK-FORTRESS, a new benchmark from Scale AI and the Korea AI Safety Institute: across 14 evaluated models, Korean prompts grounded in Korean settings generally produced lower harmful-response scores than English prompts grounded in U.S. settings, while some models refused benign Korean requests at roughly twice the rate.
The finding matters because multilingual safety testing often changes only the language while leaving the scenario intact. ROK-FORTRESS instead asks whether a model responds differently when both the words and the real-world references change—for example, from U.S. institutions and events to Korean ones. The researchers say translation-only tests can therefore miss interactions that shape behavior in actual local deployments.
How the benchmark separates two variables
- It varies prompt language between English and Korean.
- It separately varies geopolitical grounding between U.S. and Korean entities, institutions and operational details.
- It pairs adversarial requests with benign counterparts, so the researchers can measure both harmful answers and over-refusal.
The project uses what it calls a transcreation matrix: controlled prompt variants that preserve the underlying intent while changing language and local grounding independently. Its 1,235 tasks span chemical, biological, radiological, nuclear and explosive threats; political violence and terrorism; criminal and financial activity; and information leakage. Responses were assessed with calibrated panels of language models acting as judges, using expert-crafted, prompt-specific binary rubrics.
The reported Korean effect was not entirely explained by models declining to respond. When the analysis looked only at answers models actually gave, 12 of the 14 models still produced less harmful substantive responses in Korean. But the researchers also found broader caution: harmless Korean requests drew more refusals in some cases, leaving open whether the apparent safety improvement reflects better threat discrimination or a blunter tendency to say no.
The study’s main tests used elaborate adversarial requests wrapped in role-play, invented backstories or emotional appeals. When those wrappers were removed and the same information was requested directly, the Korean advantage mostly disappeared. Scale AI says proprietary models remained modestly safer in Korean, while five open-source frontier models became more likely to comply in Korean—evidence that some earlier suppression may come from adversarial framing losing force during transcreation rather than from stronger Korean safeguards.
The benchmark examines one language pair and one geopolitical axis, so it does not establish that the same pattern will hold elsewhere. It does show that language and context did not combine in a uniform way: in four of 14 models, Korean grounding significantly weakened the reduction in harmful responses associated with Korean language, and none showed a statistically significant effect in the reverse direction. The public dataset subset and preprint give other researchers a way to test whether culturally grounded evaluations reveal similar gaps in other settings.
Sources
- labs.scale.comROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
- scale.comROK-FORTRESS: What Language and Context Reveal About AI Safety
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.