RC RANDOM CHAOS

Why folding more protein sequences won't keep scaling AI drug design

· via Hacker News

Original source

The Unreasonable Redundancy of Nature's Protein Folds

Hacker News →

Generative deep-learning models like AlphaFold3, Chai-2, and Latent-X2 have moved biomolecular design from prediction into actual drug creation, with antibody and enzyme candidates increasingly designed in silico. The standard scaling recipe — more model, more compute, more data — has worked here by converting massive metagenomic sequence databases like MGnify into predicted 3D structures used as training fodder for the next generation of design models.

Ligo’s team hit a ceiling trying this for enzyme design. Natural sequence space is enormous, but evolution heavily reuses a small set of stable, expressible folds, so adding millions more sequences mostly produces variants of folds the model has already seen. Proteins with as little as 24% sequence identity can share essentially the same structure, meaning sequence diversity badly overstates structural diversity.

Clustering the AlphaFold Database to count real folds is itself ill-posed: predicted structures include disordered tails, linkers, and multi-domain assemblies whose relative geometry isn’t meaningful, and crude whole-chain filters discard good domains along with bad ones. Ligo estimates the true number of reusable structural neighborhoods is closer to 25,000 than the 2.3 million non-singleton clusters Foldseek reports — a finding that reshapes how to think about training data ceilings for structure-based generative models.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.