Yale-led study finds symbolic structure inside seven language models
Replacing internal representations with structured approximations left behavior largely unchanged. Targeted edits also changed which words a model treated as connected.
Loading page…
Replacing internal representations with structured approximations left behavior largely unchanged. Targeted edits also changed which words a model treated as connected.
Listen to this story
A preprint from a Yale-led team reports that seven large language models’ numerical representations can be approximated with explicit structures linking content to its role, with little change in behavior. The result offers a concrete account of how models may preserve relationships needed for language, arithmetic, logic, and code—not just individual words or values. The researchers also showed that editing one role-linked component can change a sentence interpretation, while cautioning that their evidence covers only one kind of symbolic organization.
The study tested seven large language models across language, arithmetic, logic, and coding, and also found the pattern in smaller networks trained to manipulate lists.
Replacing a model’s representation-generating process with tensor product representations left its behavior largely unchanged.
In a targeted edit, moving “clever” from the subject to the object shifted the interpretation from a clever doctor to a clever lawyer.
Researchers could change how a language model interpreted a sentence by editing the relationships encoded inside it. A Yale-led study described on October 8 found evidence of symbolic structure within seven models’ numerical representations—a finding that helps explain how systems built from lists of numbers can handle tasks involving language, arithmetic, logic and code.
Led by Yale computational linguist Tom McCoy, the research examines a gap between what language models do and how they represent information. Yale’s October 8 account describes a study available as a preprint, with experiments testing whether the models’ internal numerical patterns could be closely approximated by explicit symbolic structures.
Language models process information using vectors: long lists of numbers. The team substituted those representations with tensor product representations, mathematical constructions that combine a piece of content, called a “filler,” with the “role” it occupies. The distinction separates what something is from where it belongs in a structure.
The researchers replaced a model’s entire representation-generating process with these role-filler approximations. They reported that its behavior remained largely unchanged.
The team also made targeted edits rather than replacing the whole representation. McCoy described an input reading “the clever doctor helped a lawyer.” Researchers edited the vector component encoding “clever” as a description of the subject, moving it to describe the object instead. The model then behaved as though it had received “the doctor helped a clever lawyer.”
McCoy said that ability to alter behavior shows the models’ internal operations rely on symbolic structure. He limited the conclusion to a particular kind of role-filler organization, adding that the work “only scratches the surface” of understanding how neural networks encode information.
Alongside McCoy and Linzen, the study’s authors include Paul Soulos of Microsoft and Paul Smolensky of Microsoft Research.
Loading discussion...
Join the conversation
Explain whether an internal explanation changes your judgment of an answer.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.