Two trajectories cross over each other showing areas of alignment and misalignment.
Measuring metrics: The computational literature is littered with competing methods quantifying similarity in neural population codes.
Illustration by Scott Balmer
Add us as a Preferred Source on Google

Mind over metrics: How can we tell if two brains (or AI models) are alike?

The ability to record from large populations of neurons has triggered the development of myriad methods for comparing them. But we’re still grappling with how to convert measures of likeness into a better mechanistic understanding.

By Alex Williams
17 August 2026 | 7 min read

Biologists have come up with many creative strategies to understand how organisms function. Comparative analysis is among the most fundamental. Indeed, Darwin solidified his theory of evolution by comparing diverse species and arguing that their differences reflect adaptive modification over time. Today, this reasoning is so central to our thinking that entire fields of study are built on the principle of comparison.

Comparative analysis is not mere stamp collecting—it helps biologists build mechanistic understanding. For example, scientists in the 1960s found a tight correlation between the thickness of the renal medulla (the kidney’s inner region) and a species’ ability to produce concentrated urine. These comparisons, particularly across desert and nondesert mammals, were a key clue that confirmed and refined the countercurrent multiplication mechanism that underlies our modern understanding of kidney function.

What can we, as neuroscientists, learn from such examples? In many ways, comparative analysis already deeply affects our work. As a field, we lean heavily on neuroanatomical atlases that identify homologous brain structures across diverse species. And, much as in the kidney example cited above, we can even draw some correlations. For example, the size of the hippocampus correlates with the spatial navigation ability of a species.

Today, there is a rapidly growing appetite for new forms of comparative analysis between large populations of co-recorded neurons. For example, if we record from the same brain region across two animals, how can we tell whether the neural responses are the same, or different? Though such comparisons have long been possible in small invertebrate circuits, efforts in mammalian cortical systems have historically faced major technical hurdles—most notably, one could not record enough neurons in individual animals to garner the statistical power required to draw proper comparisons. The field is increasingly overcoming these obstacles as recording technologies become cheaper, miniaturized and standardized.

Furthermore, the arrival of modern artificial intelligence adds a whole new set of systems to the mix. These in silico models bear some rough resemblance to biological systems, such as the distributed nature of their computation across large ensembles of simple units. But these points of similarity are far outnumbered by differences, such as spike-based versus analog modes of communication. This naturally raises the question of whether biological and artificial networks follow similar algorithmic or computational principles in spite of their implementation-level differences. Accordingly, we have seen high-profile efforts to compare the two, such as the Brain-Score benchmark and the Algonauts Project.

In short, we now have the technical prowess to record from many different mammalian cortical systems and to manufacture, observe and manipulate powerful in silico analogues. But we are still grappling with how to link these diverse datasets together through the principle of comparison. We are still searching for answers to deeper questions, such as: What does it mean for two neural systems to be alike, and how can we rigorously quantify this likeness? Furthermore, how do we convert measures of likeness into better mechanistic understanding?

I

 acknowledge that it is difficult to come up with singular and precise answers to these questions, but we should make a concerted effort to converge on a set of core principles.

Indeed, the computational literature is now hopelessly replete with competing methods that quantify some form of similarity in neural population codes. One cluster of methods frames the problem through the lens of geometry, asking whether two systems arrange their responses in the same shape. This includes the framework of representational similarity analysis (RSA), as well as linear centered kernel alignment (CKA), which has become the de facto standard in the machine-learning research community. Others favor prediction, gauging similarity by how well the activity of one system can be used to predict that of the other. This perspective is prevalent in initiatives, such as Brain-Score, that use regularized linear regression performance as a metric of similarity. Some approaches, such as Procrustes shape distance, combine elements of both geometric similarity and prediction.

The summary above is highly incomplete—a recent review of the literature documented well over 30 methods in use. This proliferation of approaches gives us a deep well to draw from, but it also represents a serious concern. Many neuroscience practitioners—even those with computational and mathematical backgrounds—simply do not have the time to sift through this complex literature and understand its nuances. A skeptic may even feel that we are overcomplicating the problem.

Comparative measures: Metrics can help compare population activity among animals, brain regions and trials (a, b and c). Plotting activity in N-dimensional space (d) enables comparison of geometric shapes (e) and development of metric spaces (g).
Barbosa et al. bioRxiv, 2025.

Earlier this spring, I led a tutorial at COSYNE meant to make this landscape more approachable. Preparing it forced me to step back and look at the big picture. Four points have stuck with me since.

First, many popular similarity measures are closely related—more so than most people realize. In some cases, they are even essentially identical. RSA and CKA, for instance, are routinely treated as separate tools, yet they are formally equivalent once RSA is modified to include a mean-centering step. This is just one example of a broader pattern. Dig just a little bit below the surface, and a surprising number of approaches collapse onto a few underlying objects. Understanding this is extremely helpful when mentally navigating the literature.

Second, it is important not to conflate the predictive accuracy of a model with its similarity to the brain. Predictivity scores are asymmetric: Neural activity in an artificial network may be highly predictive of biological recordings, but not vice versa. Geometric measures such as RSA, CKA and Procrustes, in contrast, are symmetric. Neither framing is wrong, but they answer different questions. A high regression score says that one system carries the information needed to reconstruct the other; a high geometric score says that two systems organize that information the same way. Conflating the two invites confusion.

Third, the most versatile measures are not merely scores but proper metrics—i.e. distances that are symmetric and obey the triangle inequality, meaning that two systems can’t appear close to a third yet far apart from each other. Such metrics can be inspired by both geometric and predictive approaches. The distinction between similarity measures and proper metrics sounds pedantic, but it is the difference between a number and a map: When a measure is a true metric, the whole collection of systems becomes a space that we can navigate in a coherent fashion. We can embed brain regions and networks into a common space, cluster them and hand them to other standard machine-learning tools.

Finally, brains are complex organs—I think it is too much to ask for a single metric to quantify similarity across experiments or between a model and a biological recording. Neuroscientists should report multiple metrics to capture complementary aspects of neural computation. This requires digging into the mathematical details and assumptions of each method, which is hard work, but those who do the work will be rewarded.

This last point is perhaps the most important, but also the most challenging. It cuts against two habits the field has grown comfortable with: ranking models on a single leaderboard, and minting new metrics that are technically novel but only marginally different from the ones we already have. Both of these perspectives venerate the scores themselves over scientific understanding. In comparative analyses of the kidney, medullary thickness mattered only because it pointed toward countercurrent multiplication. Likewise, neural similarity scores matter only insofar as they point us toward computational mechanisms.

Despite these challenges, I think we stand to benefit enormously from engaging with these questions. I hope that we continue to refine and unify our understanding of existing similarity metrics, while also developing new metrics that capture genuinely new aspects of neural computation that are currently overlooked.

Alex Williams has an appointment at the Flatiron Institute, which is part of the Simons Foundation, The Transmitter’s parent organization.

AI use disclosure:

The author conceptualized and drafted the piece and consulted Anthropic's Claude for editorial suggestions before submission.

Sign up for our weekly newsletter.

Catch up on what you missed from our recent coverage, and get breaking news alerts.