Advancing AI for biology: Teaching models to design and characterize antibodies
Amazon Bio Discovery, described as an AI-powered application giving scientists access to biological AI models and integrated lab services, has published three papers on AI-assisted antibody design and characterization. Two are peer-reviewed journal papers on prediction; the third applies prediction within an end-to-end design process and reports experimentally validated antibody hits against a novel cancer target.
The first paper, "A systematic evaluation framework for universal antibody-antigen binding affinity prediction and candidate recommendation," published in iScience, introduces MochiBind, a sequence-based predictor. Rather than predicting absolute binding affinity, MochiBind predicts which of two antibodies against the same antigen binds more tightly, embedding antibody-antigen complexes with the ESM-2 protein language model and aggregating pairwise comparisons into a global ranking using TrueSkill. No structural input is required. Evaluated on the AlphaBind dataset across four antigen systems (TIGIT, PD-1, HER2, and the SARS-CoV-1 RBD), MochiBind achieved higher pairwise accuracy than every structure-based baseline on all four held-out antigens, outperforming the closest competitor by almost 10% on average. It scored 200,000 antibody pairs in roughly 13 seconds on a CPU, described as a more than 100-fold inference speedup.
The second paper, "Context-aware multi-property antibody predictor: A novel framework integrating text and protein language models," in npj Systems Biology and Applications, presents CA-MAP, which addresses batch effects during inference by taking a prompt of example antibodies with measured properties. Its AB-context-aware training strategy applies a hidden random transformation to context properties and expected answers. In hydrophobicity prediction on a fine-tuned TxGemma model, the context-aware approach retained a 0.99 Spearman correlation under a simulated additive batch effect in the 0–0.3 range, where standard fine-tuning fell to 0.58. CA-MAP uses roughly 182,000 trainable parameters versus TxGemma's 40 million and is about 200 times as fast per prompt. Trained on a synthetic dataset of 876,898 antibody-heavy chains covering six developability properties, it reached a Spearman correlation above 0.8 on several properties.
The third paper, "Agent-guided de novo design of nanobody binders against a novel cancer target," presented as a Spotlight at the ICML 2026 Workshop on Generative and Agentic AI for Biology, where it received the Best Paper Runner-Up Award, targeted a cell surface antigen for desmoplastic small round-cell tumors, a rare and aggressive pediatric cancer, identified by collaborators at the Dr. Nai-Kong V. Cheung's Lab at Memorial Sloan Kettering Cancer Center. The target has no experimental structure and no public antibody information. A hotspot recommendation agent proposed eight hotspot regions; three generative models (RFantibody, IgGM, and mBER) each produced 96,000 designs, and a selection agent prioritized 100,000 candidates for yeast-display screening. None of the 116 surviving candidates bound an unrelated control protein, and 46 were identified as strong binders.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Three new papers from Amazon Bio Discovery address bottlenecks in AI-driven antibody engineering, from benchmarking binding predictors to experimentally validating de novo design.