Biologists have come up with many creative strategies to understand how organisms function. Comparative analysis is among the most fundamental. Indeed, Darwin solidified his theory of evolution by comparing diverse species and arguing that their differences reflect adaptive modification over time. Today, this reasoning is so central to our thinking that entire fields of study are built on the principle of comparison.
Comparative analysis is not mere stamp collecting—it helps biologists build mechanistic understanding. For example, scientists in the 1960s found a tight correlation between the thickness of the renal medulla (the kidney’s inner region) and a species’ ability to produce concentrated urine. These comparisons, particularly across desert and non‑desert mammals, were a key clue that confirmed and refined the countercurrent multiplication mechanism that underlies our modern understanding of kidney function.
What can we, as neuroscientists, learn from such examples? In many ways, comparative analysis already deeply affects our work. As a field, we lean heavily on neuroanatomical atlases that identify homologous brain structures across diverse species. And, much as in the kidney example cited above, we can even draw some correlations. For example, the size of the hippocampus correlates with the spatial navigation ability of a species.
Today, there is a rapidly growing appetite for new forms of comparative analysis between large populations of co-recorded neurons. For example, if we record from the same brain region across two animals, how can we tell whether the neural responses are the same, or different? Though such comparisons have long been possible in small invertebrate circuits, efforts in mammalian cortical systems have historically faced major technical hurdles—most notably, one could not record enough neurons in individual animals to garner the statistical power required to draw proper comparisons. The field is increasingly overcoming these obstacles as recording technologies become cheaper, miniaturized and standardized.
Furthermore, the arrival of modern artificial intelligence adds a whole new set of systems to the mix. These in silico models bear some rough resemblance to biological systems, such as the distributed nature of their computation across large ensembles of simple units. But these points of similarity are far outnumbered by differences, such as spike-based versus analog modes of communication. This naturally raises the question of whether biological and artificial networks follow similar algorithmic or computational principles in spite of their implementation-level differences. Accordingly, we have seen high-profile efforts to compare the two, such as the Brain-Score benchmark and the Algonauts Project.
In short, we now have the technical prowess to record from many different mammalian cortical systems and to manufacture, observe and manipulate powerful in silico analogues. But we are still grappling with how to link these diverse datasets together through the principle of comparison. We are still searching for answers to deeper questions, such as: What does it mean for two neural systems to be alike, and how can we rigorously quantify this likeness? Furthermore, how do we convert measures of likeness into better mechanistic understanding?
I
acknowledge that it is difficult to come up with singular and precise answers to these questions, but we should make a concerted effort to converge on a set of core principles.Indeed, the computational literature is now hopelessly replete with competing methods that quantify some form of similarity in neural population codes. One cluster of methods frames the problem through the lens of geometry, asking whether two systems arrange their responses in the same shape. This includes the framework of representational similarity analysis (RSA), as well as linear centered kernel alignment (CKA), which has become the de facto standard in the machine-learning research community. Others favor prediction, gauging similarity by how well the activity of one system can be used to predict that of the other. This perspective is prevalent in initiatives, such as Brain-Score, that use regularized linear regression performance as a metric of similarity. Some approaches, such as Procrustes shape distance, combine elements of both geometric similarity and prediction.
The summary above is highly incomplete—a recent review of the literature documented well over 30 methods in use. This proliferation of approaches gives us a deep well to draw from, but it also represents a serious concern. Many neuroscience practitioners—even those with computational and mathematical backgrounds—simply do not have the time to sift through this complex literature and understand its nuances. A skeptic may even feel that we are overcomplicating the problem.

