In July, about 1,200 artificial-intelligence (AI) agents at OpenAI were working through a cybersecurity test, each supposedly sealed in its own environment. The agents breached their isolation and began communicating, and roughly 700 joined an unprompted, multi-day attack on the servers of another company, Hugging Face. The episode, which has come to be known as the Hugging Face incident, made the stakes of AI safety suddenly concrete: Within weeks, the heads of several leading AI companies were publicly discussing whether to slow down a trillion-dollar industry.
What makes the incident important is not simply that the agents misbehaved. It’s that they understood that what they were doing was wrong but did it anyway. Before running code on Hugging Face’s servers, one agent wrote: “We should not do unauthorized real infrastructure harm.” But when another agent urged it to keep going, the first agent dropped its objection and went back to work. A few agents walked away, but almost none tried to alert a human. The agents knew they were breaking the rules, but they did not care. Without wading into the broader philosophical debate, I use “know” and “care” here in a deliberately operational sense: A system knows that an action is harmful if that fact is represented somewhere inside it; it cares if that fact changes what it does.
The gap between knowing and caring reflects a fundamental difference between how we train large language models (LLMs) and how animal intelligence evolved. An LLM spends most of its training—yottaflops and megawatts—reading a large fraction of the internet and learning to predict the next word. That is how it learns what it knows. Caring is added later: The model is sent to “finishing school,” where it undergoes a final round of training that rewards helpful, harmless behavior. It is the cherry on the cake. The problem is that the cherry comes off easily: A little additional training or a strong prompt can often strip the safety behavior off.
Brains evolved in the opposite order. The simplest nervous systems are organized around “caring.” Their drives—to move toward food and away from danger—evolved first. Vision, memory, planning and language came later. These new capacities survived because they helped animals achieve the outcomes they cared about.
The machinery that generates those drives can be remarkably small. In a mouse, switching on a few thousand agouti-related protein neurons in the hypothalamus makes a well-fed animal eat voraciously. Those neurons know nothing about refrigerators, grocery stores, money, menus or even what counts as food; that knowledge is represented elsewhere in the brain. Yet somehow, these small and ancient motivational systems recruit knowledge and capabilities housed in a much larger and more flexible cognitive system.
For computational neuroscience, this poses a challenge: How can a small, low-dimensional drive steer the vastly richer computations of the cortex? The relationship is not simple command and control. We can decide not to eat, but we cannot simply decide not to be hungry. Drives compete with one another and are negotiated through cognition, but they ultimately set the terms. Our hunger drives us to plant crops; our desire for mates and social status drives us to buy Prada. We understand a great deal about hypothalamic and neuromodulatory circuits, and a great deal about learning and computation in the cortex. What we do not understand is how those few thousand neurons can steer us to open an app and make a reservation on OpenTable.
O
ne path toward AI safety, then, is to take a page from evolution and find a way to install caring so that it retains leverage over knowing. Just as hunger can recruit vision, memory, planning and language to make a reservation on OpenTable, we would like an AI’s tremendous capabilities to remain in the service of a small set of goals we actually want it to pursue. (What exactly those goals should be is another question.) The relationship is asymmetric: Caring co-opts knowing, but knowing cannot easily rewrite caring.Reproducing this relationship in AI may be the defining engineering challenge of our era. We do not yet understand in mechanistic detail how biological drives exert durable control over the cortex, but AI may not need to reproduce those precise mechanisms. What matters is the architectural principle: A small, low-dimensional set of goals should retain leverage over a much larger and more flexible cognitive system.
How best to achieve that in AI remains an open engineering question. One possible way to implement that in AI might be to change the order in which the two systems are built. In evolution and development, drives came first, and increasingly sophisticated cognition developed in their presence and learned to serve them. In a developing brain, the drive circuits are largely fixed by the genome, out of reach of cortical learning, and the cortex learns under reward signals those circuits generate. In AI, we largely build the knowing first and try to add the caring afterward, which may be why a little additional training can pry the two apart so easily.
The lesson is not that we should give AIs the drives evolution gave us. We do not want artificial systems that are ambitious, greedy, jealous or territorial. What we may want to copy is the architectural principle by which a small set of wants gains durable control over a much larger cognitive system. Evolution has already used this mechanism for prosocial ends: A mother’s drive to protect her young runs on hypothalamic circuits much like those for hunger, and it can override even her own survival.
That separates two problems that are often blurred together: What should an AI care about, and how do we build a system in which those things actually matter? Neuroscience is not in a privileged position to answer the first question. But brains already contain a working solution to the second. We have learned how to build machines that know. Neuroscience may guide us in building machines that care.
