INDEPENDENT LIFE SCIENCES BENCHMARKS

Bio-AI Model Leaderboard

Compare the scientific accuracy, structural resolution, inference latency, and licensing conditions of state-of-the-art biological foundation models across protein folding, de novo design, small-molecule docking, and single-cell genomics.

RankModel & CreatorTask DomainPrimary Accuracy BenchmarkSecondary MetricCompute FootprintLicenseBioAtlas Profile
#1AlphaFold 3★ SOTA OverallGoogle DeepMind & Isomorphic Labs · Nature (2024)Protein Structure84.2%CASP15 / lDDT-PLI1.42 ÅAll-Atom RMSDCloud API / A100 GPUNon-Commercial / ServerView Specs →
#2ESM-3Top GenerativeEvolutionaryScale · Science (2024)De Novo Design98.4%Simultaneous GFP Design1.18Sequence-Structure Perplexity8x H100 (98B) / 1x A100 (1.4B)Open Weights (1.4B) / CommercialView Specs →
#3RoseTTAFold All-AtomBaker Lab (UW IPD) · Science (2024)Protein Structure76.8%Protein-Ligand RMSD < 2Å0.89Complex TM-score1x RTX 4090 / A100Open Source (Academic Free)View Specs →
#4RFdiffusionStandard in De NovoBaker Lab (UW IPD) · Nature (2023)De Novo Design88.5%De Novo Binder Success0.94Designability TM-score1x GPU (>= 16GB VRAM)Open Source (BSD-3)View Specs →
#5DiffDockFastest DockingMIT CSAIL · ICLR (2023)Small-Molecule Docking38.2%PDBBind Top-1 RMSD < 2Å0.8s / ligandInference Speed1x GPU or CPUOpen Source (MIT)View Specs →
#6ChromaGenerate Biomedicines · Nature (2023)De Novo Design92.1%Substructure Conformity0.84Solubility ScoreCloud EnterpriseProprietaryView Specs →
#7scGPTSingle-Cell LeaderWang Lab / Stanford · Nature Methods (2024)Single-Cell / Genomics89.6%Cell Type Annotation Macro-F10.81Perturbation Prediction Pearson r1x A100 (40GB)Open Source (MIT)View Specs →
#8GeneformerBroad Institute · Nature (2023)Single-Cell / Genomics87.4%Dosage Sensitivity F10.78Chromatin Accessibility Correlation1x V100 / A100Open Source (Apache 2.0)View Specs →
#9ESMFoldMeta AI · Science (2023)Protein Structure60x faster than AF2Fast Fold Speed81.2%lDDT Accuracy1x GPUOpen Source (Apache 2.0)View Specs →

Methodology & Benchmark Integrity

Metrics are curated from published, peer-reviewed literature (Nature, Science, Cell, ICLR) and blind community evaluations (CASP15, CAMEO, PDBBind). All evaluations reflect standardized test sets without overlapping training data.

Want your model benchmarked or updated? Verified lab PIs and company representatives can submit benchmark evaluation artifacts →