Attention is all you need to fold proteins.
Every model here reads sequences or geometry, and every claim on the page traces to the paper or the model card it came from. Written for people meeting these architectures for the first time, and for everyone who has read the papers and would still like the diagram.
AlphaFold 2
Evolution in, coordinates out.
Predicts the structure of a protein from its sequence, an alignment of its homologues and optional templates.
Explore the architectureAlphaFold 3
One model for the whole complex.
Predicts the joint structure of complexes containing proteins, nucleic acids, small molecules, ions and modified residues.
Explore the architectureESM-2
Read enough proteins and structure falls out of the reading.
A bidirectional transformer trained to reconstruct masked amino acids from unaligned sequences.
Explore the architectureESMFold
Fold without searching for relatives.
Predicts a structure from a single sequence with no alignment search at query time.
Explore the architectureESM3
Sequence, structure and function as one masked prediction.
A generative masked model over three parallel tracks.
Explore the architectureProteinMPNN
Given the shape, which sequences hold it?
Proposes amino-acid sequences compatible with a supplied backbone.
Explore the architectureRFdiffusion
Denoise until a protein appears.
Generates protein backbones from design constraints by fine-tuning RoseTTAFold to denoise residue positions and orientations.
Explore the architectureESM C
The representation line, scaled further.
A transformer encoder with pre-layer-norm, rotary embeddings and SwiGLU activations, trained on sequences from UniRef, MGnify and the Joint Genome Institute clustered at 70 percent identity.
Read the entryESMFold2
A language model encoder with an all-atom diffusion decoder.
Predicts all-atom structures of proteins and their complexes from ESM C representations, with an optional alignment for difficult targets.
Read the entryLigandMPNN
Sequence design that can see the ligand.
Extends structure-conditioned sequence design to every non-protein component of a system: small molecules, nucleotides and metals.
Read the entryRFdiffusion2
Enzymes from a geometry, not from a residue numbering.
Designs scaffolds directly from the geometry of catalytic functional groups, without specifying which sequence positions those residues occupy and without inverse rotamer generation..
Read the entryProtT5 and ProtBERT
The encoders that were there first.
A family of transformers trained on UniRef and BFD with the objectives of their text counterparts.
Read the entrySaProt
One token for the residue and its local shape.
A masked language model over a structure-aware vocabulary.
Read the entryAlphaMissense
Structure-aware scoring of single amino-acid changes.
Classifies missense variants as likely benign or likely pathogenic across the proteome, adapting an AlphaFold-derived model with population frequency data rather than clinical labels..
Read the entry