LigandMPNN
Inverse folding · Baker lab · 2025

LigandMPNN

Extends structure-conditioned sequence design to every non-protein component of a system: small molecules, nucleotides and metals. It returns sidechain conformations along with sequences.

In plain language

Sequence design that can see what the protein is holding. A residue whose job is to grip a zinc ion looks arbitrary to a model that cannot see the zinc.

One way to picture it

Designing a hand without knowing what it will hold gets you a hand that grips nothing in particular.

Commonly misread as

It trusts the pose you give it. If the ligand is placed wrong in the input, the sequence will be designed around a mistake without complaining.

01 / Why it is here

Standing

Native sequence recovery at residues contacting small molecules reaches 63.3 percent, against 50.5 percent for ProteinMPNN and 50.4 percent for Rosetta. At metal-contacting residues the figures are 77.5, 40.6 and 36.0 percent.

02 / What sets it apart

Distinctions

  • Non-protein atoms enter the graph instead of being deleted from it.
  • Sidechain packing comes with the sequence, so binding can be inspected rather than assumed.
  • More than 100 designed small-molecule and DNA-binding proteins were validated experimentally, with four crystal structures.
03 / Where it stops

Limits

It inherits the assumption that the supplied backbone and ligand pose are correct. Errors upstream propagate silently into the designed sequence.

Weights and code

Checked against the registry, not from memory
RepositorySizeLicenceNote
dauparas/LigandMPNNSeveral noise levelsMIT

Sources

Each number above comes from one of these

Same task, other answers

Inverse folding