Musubi Releases an Open Moderation Model That Reads Platform Rules Without Retraining
PolicyLM-1.7B returns scores against custom policies instead of writing answers. Musubi’s tests show low latency, but also limits to how closely decisions follow policy edits.
Musubi released PolicyLM-1.7B on October 6 under an Apache-2.0 license, giving moderation teams an open-weight model that scores messages against short policy categories supplied at runtime. That can spare teams retraining for every rule edit, but does not make decisions freely programmable: the model allows up to 16 categories within a 2,048-token policy-and-message budget, and about half of tested single-clause edits changed outcomes at the balanced cutoff. On Musubi’s synthetic custom-policy benchmark, it scored 84.2% accuracy, below gpt-oss-safeguard-20B; the released model has not been tested on live traffic.
01
Teams can set score thresholds, but the model returns category scores without written explanations for its judgments.
02
Musubi’s custom-policy benchmark was also used during development, so its results are company-reported and not an independent evaluation.
03
The model handles one text message at a time, with no images or conversation history; Musubi advises against using it alone for self-harm safeguards.
Moderation teams can change their written rules without retraining Musubi’s new model, the company says—a useful promise for live conversations where decisions cannot wait for a training cycle. On October 6, Musubi announced PolicyLM-1.7B, an open-weight model that reads a policy alongside each message and returns category scores rather than a written answer.
Rules arrive with each message
The release targets live chat, game lobbies, direct messages and usernames. Musubi positions it between fixed classifiers, whose categories are set during training, and larger language models that can read policies but generate their verdicts as text.
Teams supply short, plain-language categories, including rules and exceptions, rather than paste an entire policy document. PolicyLM scores every category in one processing pass. Each score runs from zero to one; a configurable cutoff determines whether the message is flagged. It provides no written reason for that judgment.
The input budget is small: policy and message share 2,048 tokens, the chunks of text a model processes. The card allows up to 16 categories and 1,662 policy tokens, so teams must condense their rules rather than send lengthy policy manuals.
That flexibility has boundaries. Musubi says instructions adjust meanings learned during training, rather than freely redefine them—for example, turning abuse into support. The model card also warns that categories scored together can influence one another, and recommends separate calls for unrelated categories.
Policy flexibility is measurable—and uneven
On Musubi’s synthetic custom-policy benchmark, PolicyLM recorded 84.2% accuracy, below the 90.9% reported for the larger gpt-oss-safeguard-20B. Musubi says it beat every other model tested under 20 billion parameters. These are company-run results, and the custom-policy evaluation set was also used during development.
Reading an updated rule does not guarantee a changed decision. The model card says about half of tested single-clause edits changed the outcome at its balanced cutoff. Musubi advises teams to test policy changes against example messages and calibrate cutoffs on their own content before deployment.
A worked example shows the mechanism. An invented crafts-marketplace listing credits a named maker’s free fox-hat pattern. Initially, the model flags it under a copied-designs rule, scoring it 0.85. Adding an exception for credited public patterns drops the score to zero. Musubi selected the example from edits developed to demonstrate a decision change; the card calls it an illustration, not a typical result.
The 1.7-billion-parameter model starts from BidirLM-1.7B-Embedding, derived from Qwen3-1.7B-Base. Musubi trained it using public safety datasets, including NVIDIA’s Nemotron Safety Guard data, alongside synthetic and LLM-written examples. The released model can run on a laptop CPU or Apple silicon as well as a GPU.
Fast screening, not every safety judgment
Text only: it processes one message at a time, without images or conversation history.
Language coverage: Musubi evaluated messages in 19 languages, with English strongest; all tested custom policies were in English.
Enforcement limits: the model card directs child-safety cases to dedicated tools and advises against using it as the only self-harm safeguard.
The default precision preset is intended for settings where violations are rare. Balanced catches more violations when missing one matters more than a false flag. Musubi warns that benign text can still be over-flagged, including harmless requests that sound harmful and identity mentions under hate-speech rules.
Musubi says a custom fine-tuned version runs on a platform handling more than one million messages daily. That is distinct from the released model, whose card says it has not yet been tested on live traffic. The downloadable weights carry an Apache-2.0 license and can run on users’ own infrastructure.
Sources
musubilabs.aiIntroducing PolicyLM-1.7B: a small, fast, open model that reads your content policy | Musubi
Reader comments
Newest comments first. Replies stay oldest first.