Modelspublished

MIT Builds PottsMPNN to Model Protein Stability Beyond Native-Sequence Matching

The framework centers protein design on whether a sequence fits a target structure and its energy landscape, rather than whether it resembles the sequence evolution selected.

By 2 min read
MIT Builds PottsMPNN to Model Protein Stability Beyond Native-Sequence Matching

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
MIT researchers have built PottsMPNN, a protein-design system aimed at creating sequences that fit a target structure without simply copying proteins found in nature. The work, led by Foster Birnbaum and senior-authored by Amy E. Keating, shifts the design test from evolutionary resemblance to structural compatibility and stability. That matters because a protein fold is not tied to one exact sequence. Multiple amino-acid sequences can produce the same shape, while a single sequence may adopt different structures as it flexes or responds to a functional trigger. PottsMPNN therefore models what the researchers call the sequence-energy landscape: how amino acids interact at different positions, and how those combinations affect whether a protein stays stable. The system represents pairwise interactions across all 20 amino-acid choices at two positions, rather than focusing mainly on combinations already observed in nature. During training, it also introduces deliberate structural variation, or noise. The goal is to reduce direct imitation of native sequences and broaden the range of designs it can generate. MIT reports improvements in sequence generation, structural compatibility, and mutation-stability prediction, including for novel proteins. But the break with biology’s existing record is incomplete. PottsMPNN still learns from evolutionarily related sequence sets, so native-sequence information remains part of the method. The key question is whether that reduced reliance can eventually produce useful proteins that are genuinely new to nature.

Story brief

3 key points

MIT’s PottsMPNN reframes computational protein design around whether a sequence is structurally compatible and energetically stable, rather than whether it resembles evolutionarily observed proteins. The model represents all 20 amino-acid options across position pairs, adds structural variation during training, and reports better sequence generation and mutation-stability prediction, including for novel proteins. It...

  1. 01

    PottsMPNN models pairwise amino-acid interactions across all 20 residue choices at position pairs.

  2. 02

    Training noise deliberately broadens the design space and reduces direct imitation of native sequences.

  3. 03

    The reported gains include structural compatibility and energy prediction for novel proteins.

MIT researchers have developed PottsMPNN, a machine-learning framework for computational protein design intended to generate structurally feasible sequences that need not resemble native proteins. The work, led by Foster Birnbaum and senior-authored by Amy E. Keating, was described in a PNAS paper.

A protein fold is not a one-sequence problem

Many protein-design methods start with a target structure, then generate amino-acid sequences that could adopt it. But different sequences can fold into the same structure, and one sequence can potentially take different structures depending on flexibility or a functional trigger. Keating says reproducing the sequence evolution selected is therefore not the best measure of design success.

PottsMPNN instead models the sequence-energy landscape: the relationship between amino-acid identities at protein positions and a protein’s stability. In plain language, that modeling is meant to identify sequences more likely to fold into the intended stable structure. The framework incorporates physical principles governing protein structure and stability, and MIT says it improves sequence generation and predictions of how mutations affect stability.

Training for interactions, not imitation

The system uses pairwise distributions to represent interactions between amino acids at pairs of positions. The team says accounting for all 20 possible amino-acid choices at a pair of positions helps it model the sequence-energy landscape more accurately than other methods. It also adds deliberate structural variation, or noise, during training to reduce imitation of native sequences and expand the structures for which it can generate sequences.

The departure from biology’s existing record is not absolute. PottsMPNN also trains on sets of evolutionarily related sequences, teaching it that different sequences can reach the same folded structure. Birnbaum acknowledged that this remains, in some ways, a reliance on native-sequence information.

Five protein structures with amino-acid sequences, including one natural structure and four associated with designed sequences.
MIT’s illustration pairs one natural protein structure with structures associated with designed sequences. Source: news.mit.edu.

Toward proteins that do not occur in nature

The researchers report that reducing reliance on native sequences improved structural compatibility and energy prediction, including for novel proteins. Keating said the longer-term aim is to move toward useful new-to-nature proteins for diverse applications.

Sources

  1. news.mit.eduLooking beyond natural sequences