How a New Protein-Folding AI Generates Over One Billion Structures for Research

Discover how the open-source ESMFold2 tool expands structural biology data to accelerate antibody design

Written byMichelle Gaulin
| 3 min read
Representation of protein complex predicted by Protein-Folding AI
Register for free to listen to this article
Listen with Speechify
0:00
3:00

A newly released artificial-intelligence tool has dramatically expanded the catalog of predicted protein structures available to researchers. Developed by researchers at the Chan Zuckerberg Initiative Biohub in San Francisco, California, the tool has generated an open-source database known as the ESM Atlas. This database contains more than one billion predicted protein structures and billions more protein sequences, providing an unprecedented resource for biological research across the globe.

The new database eclipses the widely used AlphaFold Database by more than 800 million entries and surpasses a previous version of the ESM Atlas by roughly 300 million entries. The predictions were generated using ESMFold2, an artificial-intelligence model that its developers report outperforms or matches existing structure prediction systems, including AlphaFold3, on several benchmark tasks.

How the new protein-folding AI expands the structural database

The ESMFold2 model is based on a protein language model unveiled in 2024, which researchers trained on billions of proteins from across the tree of life. The model incorporates extensive metagenomic sequences from soil, oceans, and other environments, substantially expanding the representation of environmental proteins beyond those included in widely used structural databases.  

According to Alex Rives, PhD, the science head at Biohub who led the effort, the project sheds light on the most uncharacterized elements of biology. "What this atlas does is it shows the totality of protein biology and especially the parts that are most unknown," Rives says. "We think it's going to be a really powerful substrate for the discovery of new biology."

The resulting atlas contains 1.1 billion predicted protein structures alongside sequence data for 6.8 billion proteins. Using this expansive resource, researchers have already mapped structural similarities between microbial defense mechanisms and gene-editing proteins found in a soil fungus. Computational biologists like Gemma Atkinson, PhD, at Lund University in Sweden, call the atlas an extraordinary resource for understanding how large-scale language models capture fundamental biological rules.

Designing better antibodies and complexes

The development team reports that ESMFold2 outperforms existing tools in determining the correct structure of protein complexes—including antibody molecules binding to their antigen targets. In the accompanying preprint, the researchers demonstrated how they used ESMFold2 to design new antibodies and proteins that bind strongly to targets implicated in cancers and immunological conditions. When synthesized and tested experimentally, a high proportion of the designed proteins performed as predicted by the model.

Other computational biologists view the resource as a major step forward. Christine Orengo, PhD, a computational biologist at University College London, notes that these predictions could help uncover entirely new protein folds and functions. While some experts, like Martin Steinegger, PhD, of Seoul National University, question how well the model handles highly unusual structures, others see immediate utility. Sergey Ovchinnikov, PhD, a computational biologist at the Massachusetts Institute of Technology, views the ESM Atlas as a valuable supplement to the 200 million structures in the AlphaFold database. Because ESMFold2 is fully open-source with no commercial restrictions, researchers note that it may lower barriers to adoption across academic and commercial research environments.

Operational impacts of open-source tools on research workflows

For laboratory leaders involved in computational biology and therapeutic discovery, the availability of a large open-source structural database may reduce barriers to accessing protein structure predictions. Researchers could potentially use these predictions to prioritize experimental candidates before undertaking more resource-intensive laboratory validation. However, predicted structures still require experimental confirmation, particularly for novel or highly unusual proteins.

Early results suggest ESMFold2 could support antibody discovery and protein design workflows by helping researchers identify promising candidates for further testing. For laboratories engaged in computational biology, the availability of a large open-source structural database may provide new opportunities to explore proteins from previously underrepresented environmental sources. As with other prediction tools, however, experimental validation remains necessary to confirm biological function and therapeutic potential.

This article was created with the assistance of Generative AI and has undergone editorial review before publishing.

Add Lab Manager as a preferred source on Google

Add Lab Manager as a preferred Google source to see more of our trusted coverage.

Frequently Asked Questions (FAQs)

  • What is the ESM Atlas?

    The ESM Atlas is an open-source database developed by the Chan Zuckerberg Initiative Biohub that contains over one billion predicted protein structures and billions of protein sequences, significantly expanding the resources available for biological research.

  • How does ESMFold2 compare to AlphaFold?

    ESMFold2 outperforms AlphaFold by generating predictions that surpass the AlphaFold Database by more than 800 million entries and is reported to match or outperform existing structure prediction systems on various benchmarks.

  • What are the applications of ESMFold2 in research?

    ESMFold2 can assist in designing better antibodies and proteins, help in therapeutic discovery by enabling researchers to identify promising experimental candidates, and explore proteins from underrepresented environmental sources.

  • Is the ESMFold2 model open-source?

    Yes, ESMFold2 is fully open-source and free from commercial restrictions, making it accessible for both academic and commercial researchers.

  • What are the limitations of predicted protein structures?

    Despite the advancements, predicted protein structures still require experimental validation, particularly for novel or highly unusual proteins, to confirm their biological function and therapeutic potential.

About the Author

  • Headshot photo of Michelle Gaulin

    Michelle Gaulin is an associate editor for Lab Manager. She holds a bachelor of journalism degree from Toronto Metropolitan University in Toronto, Ontario, Canada, and has two decades of experience in editorial writing, content creation, and brand storytelling. In her role, she contributes to the production of the magazine’s print and online content, collaborates with industry experts, and works closely with freelance writers to deliver high-quality, engaging material.

    Her professional background spans multiple industries, including automotive, travel, finance, publishing, and technology. She specializes in simplifying complex topics and crafting compelling narratives that connect with both B2B and B2C audiences.

    In her spare time, Michelle enjoys outdoor activities and cherishes time with her daughter. She can be reached at mgaulin@labmanager.com.

    View Full Profile

Related Topics

Loading Next Article...
Loading Next Article...
Current Magazine Issue Background Image

CURRENT ISSUE - May/June 2026

The ROI of Actionable Data

Break Down Silos by Ensuring Data Flows Seamlessly Between Instruments and Analytics Tools

Lab Manager May/June 2026 Cover Image