Researchers retrofit AI models to read bytes using under 1% of a typical training budget
The Nature study reuses existing models rather than rebuilding them. Its reported gains include character-level reasoning and some coding tasks, not a universal improvement across every benchmark.
A Nature study published October 7 introduces byteification, a way to retrofit pretrained language models to process text as bytes rather than fixed word fragments. The method was tested on four model families and uses less than 1% of a typical pretraining budget, excluding the cost of creating the original models. Results suggest byte-level processing can preserve much of a model’s existing capability while improving selected character-sensitive and coding tasks, offering an alternative to training byte-level models from scratch—not proof that byteified models are universally better.
01
The two-stage conversion used 9.8 billion tokens and 39.3 billion tokens, totaling 49.1 billion; source-model training is not included.
02
Researchers converted models based on Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base, and Llama 3 8B.
03
Bolmo 7B scored 16.5 percentage points above scratch-trained BLT 7B on STEM tasks, a comparison between byte-level models.
Existing language models can be converted to process the bytes underlying text without starting training over. In a Nature study published October 7, researchers describe “byteification,” a retrofit using less than 1% of a typical pretraining budget. They report that the converted models approach their originals’ capabilities while gaining stronger character-level understanding.
The work comes from researchers at LMU Munich, the Allen Institute for AI, Cambridge, the University of Washington and Imperial College London. Their approach connects two paths that have largely developed separately: improving conventional language models and building models that read text at a finer level.
Replacing fixed word fragments
Most language models split text into tokens—words or word fragments—before processing it. That makes the input manageable, but obscures the individual characters within each chunk. The paper identifies this as a particular problem for code and biological sequences, where meaning can depend on fine details rather than familiar words.
Byte-level models instead operate on text’s underlying computer encoding. Earlier approaches generally trained these models from scratch. The authors argue that this leaves byte-level research struggling to keep pace with improvements in conventional models’ training data, architecture and post-training—the work used to adapt a pretrained model.
Byteification keeps a large model at the center, with smaller components handling bytes around it. A local encoder builds representations of the input bytes. A learned boundary predictor groups them into “patches,” each containing one or more bytes. Those patches pass through the large model before a local decoder turns the information back into next-byte predictions.
A key design change concerns where those patch boundaries fall. Earlier byte-level architectures used only past context to place them. This method uses one byte of future context when processing the input, more closely matching conventional tokenizers. Those tokenizers also inspect later characters when deciding whether text belongs in one chunk or several.
The conversion’s training budget
49.1 billion tokensTwo-stage conversion
The paper reports 9.8 billion tokens in the first stage and 39.3 billion in the second.
Four conversions, not one special case
The researchers applied the method to several existing models, rather than demonstrating it on a single source architecture. The resulting names identify their origins:
Bolmo 7B starts from Olmo 3 7B.
Bolmo 1B starts from OLMo 2 1B.
Bwen 8B starts from Qwen3 8B Base.
Blama 8B starts from Llama 3 8B.
The small budget is for the retrofit, not the training already invested in those source models. Reuse is central to the method: the paper also shows that existing components from a source model’s ecosystem can adapt a byteified model without extra training cost. The authors present conversion as complementary to training byte-level models from scratch, not a replacement for that research.
Where the gains show up
In the paper’s benchmark results, the converted models outperformed earlier publicly available byte-level models of comparable size on average. Bolmo 7B scored 16.5 percentage points higher on STEM tasks than BLT 7B, which was trained from scratch. That comparison is against another byte-level model, not a claim of the same improvement over Bolmo’s conventional source.
Against its source Olmo 3 model, Bolmo 7B showed much stronger character understanding and better performance in certain coding settings. Bwen 8B outperformed Bolmo 7B, approaching and sometimes exceeding its source Qwen model. The results support a narrower conclusion than universal superiority: much of the original capability survives, with gains in particular tasks.
Valentin Hofmann, the study’s last author, points to spelling a word backward as one task where byteified models are substantially better. In LMU’s account of the research, he links fine-grained text representation to work with code and biological sequences. The models, code and training data are publicly available, allowing others to examine and reproduce the work.
Sources
nature.comRetrofitting language models to operate over bytes - Nature
newswise.comArtificial Intelligence: Language Models That See Every Letter | Newswise
Reader comments
Newest comments first. Replies stay oldest first.