A newly released artificial-intelligence tool has dramatically expanded the catalog of predicted protein structures available to researchers. Developed by researchers at the Chan Zuckerberg Initiative Biohub in San Francisco, California, the tool has generated an open-source database known as the ESM Atlas. This database contains more than one billion predicted protein structures and billions more protein sequences, providing an unprecedented resource for biological research across the globe.
The new database eclipses the widely used AlphaFold Database by more than 800 million entries and surpasses a previous version of the ESM Atlas by roughly 300 million entries. The predictions were generated using ESMFold2, an artificial-intelligence model that its developers report outperforms or matches existing structure prediction systems, including AlphaFold3, on several benchmark tasks.
How the new protein-folding AI expands the structural database
The ESMFold2 model is based on a protein language model unveiled in 2024, which researchers trained on billions of proteins from across the tree of life. The model incorporates extensive metagenomic sequences from soil, oceans, and other environments, substantially expanding the representation of environmental proteins beyond those included in widely used structural databases.
According to Alex Rives, PhD, the science head at Biohub who led the effort, the project sheds light on the most uncharacterized elements of biology. "What this atlas does is it shows the totality of protein biology and especially the parts that are most unknown," Rives says. "We think it's going to be a really powerful substrate for the discovery of new biology."
The resulting atlas contains 1.1 billion predicted protein structures alongside sequence data for 6.8 billion proteins. Using this expansive resource, researchers have already mapped structural similarities between microbial defense mechanisms and gene-editing proteins found in a soil fungus. Computational biologists like Gemma Atkinson, PhD, at Lund University in Sweden, call the atlas an extraordinary resource for understanding how large-scale language models capture fundamental biological rules.
Designing better antibodies and complexes
The development team reports that ESMFold2 outperforms existing tools in determining the correct structure of protein complexes—including antibody molecules binding to their antigen targets. In the accompanying preprint, the researchers demonstrated how they used ESMFold2 to design new antibodies and proteins that bind strongly to targets implicated in cancers and immunological conditions. When synthesized and tested experimentally, a high proportion of the designed proteins performed as predicted by the model.
Other computational biologists view the resource as a major step forward. Christine Orengo, PhD, a computational biologist at University College London, notes that these predictions could help uncover entirely new protein folds and functions. While some experts, like Martin Steinegger, PhD, of Seoul National University, question how well the model handles highly unusual structures, others see immediate utility. Sergey Ovchinnikov, PhD, a computational biologist at the Massachusetts Institute of Technology, views the ESM Atlas as a valuable supplement to the 200 million structures in the AlphaFold database. Because ESMFold2 is fully open-source with no commercial restrictions, researchers note that it may lower barriers to adoption across academic and commercial research environments.
Operational impacts of open-source tools on research workflows
For laboratory leaders involved in computational biology and therapeutic discovery, the availability of a large open-source structural database may reduce barriers to accessing protein structure predictions. Researchers could potentially use these predictions to prioritize experimental candidates before undertaking more resource-intensive laboratory validation. However, predicted structures still require experimental confirmation, particularly for novel or highly unusual proteins.
Early results suggest ESMFold2 could support antibody discovery and protein design workflows by helping researchers identify promising candidates for further testing. For laboratories engaged in computational biology, the availability of a large open-source structural database may provide new opportunities to explore proteins from previously underrepresented environmental sources. As with other prediction tools, however, experimental validation remains necessary to confirm biological function and therapeutic potential.
This article was created with the assistance of Generative AI and has undergone editorial review before publishing.









