TLDRocket
Sign in

Berlin’s Apheris and Ginkgo bring pharma giants together to train AI on 10,000 antibodies

Tech.eu Cate Lawrence

Apheris and Ginkgo just launched a consortium to train AI on 10,000 antibodies. Big pharma keeps its data private, but shares enough to build better developability models.

Based on reporting by Tech.eu, Cate Lawrence — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Apheris and Ginkgo Datapoints have kicked off a new Antibody Developability Consortium with founding members from AbbVie, argenx, Lundbeck and Takeda. The goal is blunt and practical: build a standardized antibody dataset large enough to help pharma spot manufacturability and developability problems earlier, before promising candidates burn time and money.

The pitch here is not just more data, but cleaner data. Existing models have struggled because the underlying datasets are small, fragmented and inconsistent. Even when companies have sizeable internal collections, sequence diversity can still be thin. The consortium is meant to fix that by creating a purpose-built dataset and training AI models on it at scale.

Each founding member will contribute proprietary antibody sequences, while Ginkgo Datapoints fills any gaps from public sources to reach 10,000 antibodies in total. Apheris provides the federated infrastructure so members can train, benchmark and refine models without exposing raw proprietary sequences to each other. The company says members keep ownership of the sequences and assay data they contribute, while also being able to use the resulting models internally.

Ginkgo Datapoints is handling the scientific design and lab work, including sequence selection, antibody production and high-throughput characterization across core developability endpoints. It is also training a foundation antibody developability model inside Apheris’s secure environment. Independent oversight will come from Charlotte Deane of the University of Oxford and Peter Tessier of the University of Michigan.

The consortium plans to deliver its first dataset to members by early 2027. Over time, it also wants to add more complex antibody formats and other properties, which hints at a broader play: not just predicting whether an antibody can be made, but deciding earlier which drug candidates should move forward at all.

My take — AI-written commentary, not fact-checked reporting

This is the rare pharma AI story that sounds less like hype and more like plumbing, which is why it matters. The clever bit is not the model; it’s getting rivals to pool enough data to make the model useful while keeping the raw sequences locked away. That’s the kind of boring federation that could actually move the field, which is probably why it won’t get the loudest headlines.

Read more about this at: Tech.eu

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.