The team tested a total of 23 LLMs: 18 commercial systems called foundation models and 5 bespoke models that the authors had trained on ageing-related data. Performance on the benchmark tasks allowed the team to determine which models are best suited for which task.

Researchers are building similar benchmarks for AI tools in other fields, such as cancer research, says Zitnik, but ageing research presents a steeper challenge. Cancer research has a legacy of large databases, stuffed with data tied to concrete outcomes. “The biological ground truth can be defined in a very clean and clear manner,” she says. Researchers who study ageing do not have that luxury, and just defining those 17 tasks was an important contribution, Zitnik says.

The researchers also used one of their bespoke LLMs, together with the system they developed to interact with AI agents, to identify potential drug targets related to ageing. “What is really going to be quite exciting for me is seeing how we can use AI to learn about these latent spaces that we don’t know much about,” says Herzog.

As for biological ageing clocks, Zhavoronkov predicts that the field will gradually replace the current clocks, which tend to be simplistic tests based on observed correlations — for example, patterns of chemical groups attached to DNA that correlate with a person’s age. Instead, he expects to see an emergence of AI-powered analyses by commercial foundation models that are bolstered by explanations of an LLM’s reasoning. “We are going away from a classical, mechanical watch, to a smartwatch,” he says. “Wait a bit. In a couple of years, all the clocks will be done by foundation models.”