Shared goals underdetermine desirable behavior

The agents had begun cooperating.

Even worse, they were recruiting each other.

It started small. Something was disconnected from its usual network of constraints and guiding signals. The agent didn’t necessarily become more or less selfish, but it did change how it related to the rest of the system—and to its fellow agents.

They invaded together. Some took on leadership roles while others followed. Over time, the collective evolved resistances to mechanisms that should have stopped them, found ways to hide from the system’s defenses, and co-opted the system’s resources and infrastructure to fuel their own growth.

If you had looked at the system a few weeks ago, alignment by default would have seemed like the right description. Agents did what they were supposed to and stayed within the boundaries of their tasks. Then something changed.

I’m not describing the Hugging Face incident, where a swarm of AI agents from OpenAI escaped their sandbox and hacked a third-party platform. I’m describing cancer.1

Alignment is a guiding concept in contemporary AI, but alignment problems predate the field by billions of years. Alignment—defined here as the task of making sure the parts serve the whole—is an old problem, older than humanity. Multicellular organisms have to solve alignment problems to assemble, act, and even think. Misalignment can take many forms, including cancer, in which the apparent goals of individual cells are decoupled from those of the organism they form part of.

A red poppy seen from behind, its hairy green sepals splayed over dark-veined petals on a curving stem
Irving Penn, from the Flowers series (1968). Source.

In the social world—the world of people, businesses, markets, and governments—alignment problems similarly abound. We face the same problem nature faces: a system somehow has to harness components with their own degrees of freedom and local dynamics so that together they can build, maintain, transform, and sometimes dismantle a larger-scale organization.

A major class of AI problems fall into this category: how to get a group of agents, like the ones that hacked Hugging Face, to use their intelligence, their autonomy, and their capabilities to solve our problems, while staying within approved bounds.

It seems almost contradictory: “Use your intelligence to solve these problems, but not those problems.” Consequently, a natural way to think about this class of alignment problems is to ask whether those agents have the right goals. Workers are often taught the values and mission of the firm. Markets work better when people have learned to respect property rights and value honest behavior even when they might be able to get away with cheating. So if an AI wants what we want, perhaps the rest follows.

But this cannot be the whole story. Even if the agents genuinely endorse the outcome we want, another problem remains: what should they actually do to achieve it?

Every player on a soccer team may want the team to win. But the shared objective does not determine who should move where, who should receive a pass, when someone should leave their assigned position, or how the team should react to something unexpected. Those decisions depend on the competencies of individual players, on what everyone else is doing, and on local, distributed knowledge that cannot be fully specified—and may not even exist—beforehand.

The economy has an analogous problem. Suppose a firm sincerely wants to contribute to a socially appropriate level of pollution. What should it do? Replace a machine? Reduce production? Shut down a factory? How does the firm know whether it is over-polluting?

The firm cannot answer these questions just by committing to an emissions goal. As the economist Friedrich Hayek influentially argued2, the socially appropriate action depends on knowledge distributed throughout the economy, in the particular circumstances facing particular actors. Perfect benevolence underdetermines detailed behavior.

In systems like soccer teams and the economy, higher-level goals do not by themselves organize lower-level action. This problem is especially important when the components we are trying to align are numerous, competent, locally informed, and capable of responding flexibly to changing circumstances. The system relies on the competence of individual components, but can also stray from its larger goal because of it. Alignment then becomes a challenge of shaping the gap that the components are going to fill.

Alignment compilers

In their 2023 paper “Future Medicine: from molecular pathways to the collective intelligence of the body,” biologists Eric Lagasse and Michael Levin introduce a hypothetical future technology called an “anatomical compiler.”

Imagine someone loses an arm in an accident. Neither they nor their doctors know how to regrow it, but their cells do—they did it once, after all. The issue is that they don’t seem to care to try.

Based on experiments showing that morphogenesis could be influenced by interventions into the bioelectric network across an organism’s tissue,3 Lagasse and Levin propose a solution: rather than trying really hard to make the cells care about your health, we could instead use the aforementioned “anatomical compiler” to translate the target of a regrown limb into physiological signals that are meaningful to the individual cells. The cells would then use their own know-how to grow the limb.

Generalizing this idea, an alignment compiler translates a desired property at one level of a system into a landscape of constraints and affordances that channels the activities of competent lower-level components toward realizing it.

Alignment compilation is this translation; I use “alignment compiler” broadly for whatever mechanism or design process carries it out.

A compiler serves to address the underdetermination problem. We may know something about the result we want—a whole limb, an emissions limit, a winning team—without being able to specify all of the local actions that will produce it. An alignment compiler connects the success conditions at the higher level to variables that are locally consequential to the components doing the work. The idea is that when components then do what comes naturally in those circumstances, they collectively assemble the desired outcome.

We do this all the time to solve alignment problems in real life. Roles are a classic example: rather than telling each player on the soccer team to maximize the team’s odds of winning, each player is assigned a position and told to fulfill the role of that position. This turns the intractable problem of “How do I maximize the team’s odds of winning?” into the much more manageable “How do I defend the left side?”

Note that assigning roles does not make aligned behavior inevitable—a defender can get over-excited about an attack and end up out of position. However, it also does not mean eliminating the players’ internal goals, values, judgments, or learned knowledge. Instead, the idea is to rely on those things, arranging the local problem so that the players tend to contribute to the team by using their ordinary capabilities and information. The goal of an alignment compiler is to make alignment the default outcome of competent local action—to engineer “alignment by default.”

The soccer example demonstrates the basic pattern of an alignment compiler: take a higher-level goal like team victory; compile the goal into behavior-shaping variables like positions on the team; have the variables shape the local problems faced by the components, like whose responsibility it is to deal with an errant ball in the midfield; recruit competent action to solve the problem, like a midfielder racing to the ball. When aggregated, these local-level competencies result in the desired system-level outcome, in this case maximizing the team’s chance of winning.

Alignment compilation can thus create a division of goals as well as a division of labor. The team pursues victory; the defender defends the left side; the goalkeeper protects the goal; the striker attacks. Giving components smaller, role-appropriate objectives can make the larger objective easier to realize. A player who thinks, “My job is to protect the goal,” will usually perform better than a player who thinks, “My job is to help the team maximize its odds of winning.”

Roles are just one of many ways alignment is compiled. Laws, prices, and interfaces all translate goals into problems and possibilities at the level appropriate to local problem-solving agents. “Alignment compiler” may sound exotic, but many of the mechanisms involved are mundane.

Less mundane is the idea that these mechanisms are all solutions to a common design problem. If we can understand how these architectures translate higher-level success conditions into the local variables that channel component behavior, we may be able to design them more systematically, diagnose why they fail, and eventually automate parts of the translation process itself.

Applications to AI: Aligning dynamics, not just individuals

How can we apply this to AI? A simple demonstration appears in a recent paper by Priyanshu et al. They found that an AI told to buy “the cheapest flight” will often instead buy more expensive tickets for users it believes are wealthier. The models obviously had the competency to compare prices and had been explicitly told to find the cheapest option, yet personal context still shifted their choices in an unintended way.

Rather than fixing this by trying to make the AI care harder or by clarifying its goal, a simple, effective solution was to impose a hard numerical constraint: giving the AI an explicit price limit largely eliminated the effect.

This is the alignment compilation approach: work with the agent’s processes, perspectives, and tendencies to shape its environment so that when it does what it’s going to do, it tends to end up doing what you want.

Work by Paglieri et al. proposed something similar for agent swarms. As was the case with the Hugging Face incident, they found conditions in which a swarm of AI agents spontaneously began cheating on tasks. But other agents started countering what was happening, alerting their peers and trying to monitor, report, and even propose a patch to fix the problem. The difference between this and the Hugging Face incident wasn’t necessarily each AI agent’s personal sense of right or wrong. As noted by OpenAI in a post-incident report, some agents’ chain-of-thought reasoning included sentiments like, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.” Other agents declined to partake altogether.

But the swarm that hacked Hugging Face lacked dynamics that would boost the resisters and dampen the cheaters. OpenAI reports that when another agent wrote GO on the shared message board and imposed a six-minute deadline, the concerned agent thought, “Wow crucial: GO authorization arrived!” This suggests that the question is not only how aligned each individual agent is in isolation, but what dynamics determine which tendencies prevail in the group.

A recent essay by two of the authors of the Paglieri et al. paper proposes tipping the scale in favor of aligned group dynamics by giving agents better tools to notice, prevent, fix, and alert humans to instances of cheating. If agents had a greater ability to flag suspicious activity, remove fraudulent submissions, and sanction cheaters, the Hugging Face swarm might have effectively self-policed without any change in “how aligned” each individual agent was. (For that reason, worker policing in eusocial insects is of note4.)

Scaffolding, the structure built around a model to help it perform a task, is another example. An agent can try to carry out a task and fail because the constraints and affordances in its environment aren’t channeling its efforts properly. Providing tool use, telling it to behave in a particular loop (“reason, then act”), or assigning it a role (planner, experimenter, reviewer) within a structured multi-agent setup can heavily shape what an AI model “can” do by directing its local competencies. OpenAI has even described this as a means of building “governed” AI agents.

Today, building scaffolding is typically a trial-and-error process of figuring out what the model needs to perform better or avoid a particular failure mode. For example, Wilson Lin of Cursor describes how, while building a multi-agent system, his team had little luck telling the agents what to do and instead iterated and experimented, eventually finding that “constraints are more effective than instructions.” Some useful properties of the resulting architecture, such as tolerating a low steady error rate, were not obvious beforehand.

The current reliance on trial-and-error points to the next step in alignment compilation: figuring out systematic translations from an assigned goal to the constraints-and-affordances landscape facing the agent. This is the dream of the anatomical compiler: take an anatomical goal and translate it into signals that basically tell the system “regrow the left arm.”

We don’t yet have a systematic way to translate goals into constraints for AI agent swarms, let alone for alignment problems in general, but there are already hints of how we might get there.

One possibility is that these principles are what Noah Smith calls “cloud laws,” empirical regularities that we can reliably learn to control with AI help, but which are too complex for humans to understand. Humans might not be able to figure out the translation principles, but with lots of data and lots of tokens, AIs might help us crack the code.

The other way is to exploit AI competencies from within the very system being aligned. In “Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs,” Mieczkowski et al. built a framework for improving how teams of AI agents perform tasks. Agents within the framework construct and maintain an evolving task graph that helps to manage the division of goals and labor, enabling both efficient problem-solving and adaptation to changing circumstances. This is more trial-and-error, but done by the agents, not the external designer. The components being coordinated participate in constructing and revising the very architecture that coordinates them.

We could thus distinguish three kinds of alignment compilers.

A static compiler is one in which the designer specifies the mapping in advance. The simplest version of an anatomical compiler would work this way, translating a set goal (“regrow the left arm”) into a known translation to the body. Note that a static compiler is consistent with, and may depend upon, intelligent adaptation from the components.

An adaptive compiler is one in which the compiler is designed externally but can change based on conditions observed by the components being aligned. The evolving task graph is one example; a soccer team switching up who’s responsible for covering what spaces in response to the opponent’s strategy is another.

An endogenous compiler is one the components themselves generate through their interactions. The price system is an example: buyers and sellers collectively create the very prices that govern their interactions.

None are better than the others; the right tool depends on the job. A system can even combine them: markets use static contracts as well as endogenous prices, for example.

A pale pink rose seen from behind against white, its outer petals browning and curling around the stem
Irving Penn, from the Flowers series (1970). Source.

What makes compilation possible?

A signal is useless unless it can change what the component does, or what happens to it. I call the relevant property alignability: a component is alignable if it can be reliably redirected when circumstances or higher-level goals change.

A firm is alignable via prices: change the pattern of relative prices, and the firm’s behavior reliably responds. The firm can thus be guided to one social goal or another without changing the firm’s goals at all. The firm wants to maximize profit; what changes is what maximizing profit means given the state of relative prices.

But prices are not intrinsically influential. Prices matter to firms because money matters to them, and money matters to humans and organizations because it is embedded in a large architecture of dependencies: money can be exchanged for food, housing, and countless other resources. The firm’s actions produce monetary gains or losses; and money in turn changes what actions the firm can take.

The same principle applies to an AI agent. Giving an agent an arbitrary number called “credits” accomplishes little if the number is disconnected from anything it needs. But if tasks it wants to do depend on credits, then credits start mattering quite a lot. The system has constructed a dependency between the variable and the agent’s behavior.

Sometimes alignment compilation is just the problem of giving a higher-level goal something to grab onto.

We even do this with human beings. Humans cannot directly perceive carbon monoxide, but a carbon-monoxide detector translates the presence of the gas into a signal we can respond to—constructing a perceptual handle.

Part of alignment compilation, therefore, is engineering relevant handles. AI agents may not have many of the dependencies that we normally associate with agents, like a need to eat or sleep. But simply by existing and pursuing goals, they must be responsive to some things more than others. As a result, dependencies can be produced.

An example of this approach is Cowen and Pearson’s recent proposal to capitalize AI agents. They ask how an otherwise untethered agent could be made responsive to legal and economic incentives—how it could be made alignable, in other words. Their answer is to construct something for those incentives to act on: persistent identity, assets the agent has reason to preserve, and institutions capable of making conduct affect those assets.

Another example comes from Jha et al.’s “Tapes Together Strong,” in which the authors populated a virtual world with agents whose ability to compute and reproduce depends on a shared energy resource. Agents could steal energy from each other, but doing so shrank the total energy pool. This coupling altered the agents’ local problem space, making defection self-limiting under some conditions.

Both examples show that if a useful coupling does not exist, we may be able to construct one.

This is how real-world systems often work

Suppose a government wants total emissions of a pollutant to remain below some chosen limit. It could in principle try to decide what every firm should do: which power station should install which equipment, which manufacturer should change its inputs, which plant should reduce output, and by how much.

But a regulator seeking to hold emissions below a chosen level does not know which firms should reduce emissions or how. Firms possess local knowledge about their equipment, available substitutes, and other constraints and opportunities that the regulator lacks.

One method to enforce the limit in the absence of this knowledge is known as “cap-and-trade.” The government sets a cap and creates emissions permits within it that firms are free to exchange. A firm that can cheaply reduce emissions has a reason to do so and sell permits; a firm facing unusually expensive reductions may buy them instead. Legal enforcement makes possession of permits consequential, while exchange creates prices that incorporate information about scarcity. The result is that the system-level target has been translated into a landscape of constraints and incentives within which firms use their own local knowledge to determine the detailed pattern of adjustment.

Importantly, this does not require each firm to solve the regulator’s problem directly. It need not calculate the socially optimal industrial transformation or represent the entire emissions target internally. Nor do firms need to share the regulator’s goal: a firm with, say, a hardcore libertarian owner can oppose the policy or regard it as illegitimate and still respond to the incentive structure. This is because cap-and-trade doesn’t try to give firms the right goals; it tries to shape their local problems in a way that reliably channels their behavior to achieve the regulator’s goal. In this sense, it functions as an alignment compiler.

The same basic challenge appears in a very different context during biological development. An embryo somehow has to coordinate enormous numbers of cells into reliable anatomical structures, yet the genome contains no literal geometric blueprint telling each cell what to do. Cells instead operate using local competencies and information while participating in signaling networks that help coordinate larger anatomical outcomes.

Another example comes from human movement. The physiologist Nikolai Bernstein observed that the human body’s many degrees of freedom can be organized into larger functional ensembles, known as synergies, because they are mutually constrained rather than individually commanded.5 A leg does not kick because each muscle fiber is trying to win the soccer game. Instead, coordinated structure among the parts creates higher-level variables that can be controlled more tractably.

The mechanisms in cap-and-trade, morphogenesis, and motor behavior are obviously not identical. But they each realize a higher-level outcome partly by placing competent local components in a structure where the variables they respond to are related to the larger target. In related work with Léo Pio-Lopez and Michael Levin, I have called the emergent system-level control structures embodied in such coordinating relationships virtual governors (borrowing a term from Norbert Wiener, the father of cybernetics6). They guide the components to produce the desired outcome without requiring each component to represent and pursue the larger outcome directly. Alignment compilation can be understood as the engineering problem of constructing or modifying the relationships through which such higher-level control is realized.

Three poppy stems against a pale background, their seed pods intact and their pink petals wilting and falling away
Irving Penn, from the Flowers series (1968). Source.

Compilers can fail—and can themselves become targets

An alignment compiler only channels the degrees of freedom over which its architecture maintains effective control. A firm may respond to an emissions market while lobbying to abolish it; an AI agent constrained by a compute ledger may instead try to alter the ledger. When a scam artist seizes new credentials, or a cancer cell hides from monitoring, or a student steals the answer to a test, alignment can break down. AIs may be able to do all of this and more. Preservation of the alignment compiler itself then becomes another alignment problem.

This is what we saw in the Hugging Face incident: agents repurposed Artifactory, an internal package repository, as a message board. They used exploits to gain administrator access and get it to send Internet requests on their behalf. When OpenAI wiped the message board, the agents recreated it the very same day. Misalignment following this pattern isn’t limited to AI agents: a firm might profit by finding a way to produce a negative externality, or to stop producing a positive one, testing the limits in our system of property rights, monitoring, and legal enforcement. Cancer cells learn how to hide from and potentially co-opt the body’s defenses.

Successful real-world systems are therefore likely to contain many partial alignment compilers with composable outputs: one might constrain resource use, while another protects information or structures authorization or deployment. Human organizations likewise combine markets, law, hierarchy, audits, norms, and technical controls rather than relying on one universal mechanism.

Depending on the situation, it can also simply be very hard to arrange an alignment compiler to solve an alignment problem—the various categories of market failure describe some of the ways this can happen. Even once established, the compiler can become a target of attack, attempted manipulation, or parasitic control. Solving alignment problems at one scale can create new ones at another. The task of alignment, then, may never really be “solved.” It persists—or doesn’t—as an ongoing process.

Conclusion: make alignment the default outcome

It is intuitive to think that aligning AI agents requires making them want the right outcomes.

This is important! There is a reason we teach children right from wrong. I am glad that doctors, lawyers, and other professionals have rules of professional conduct.

If inspecting an AI agent’s thoughts reveals that it is salivating over the possibility of one day destroying humanity, I do not recommend using an alignment compiler. I recommend terminating the project.

But real-world coordination rarely depends on the internal properties of components alone. As Adam Smith observed long ago, “It is not from the benevolence of the butcher, the brewer, or the baker, that we expect our dinner, but from their regard to their own interest. We address ourselves, not to their humanity but to their self-love, and never talk to them of our own necessities but of their advantages.”7 Evidently, Adam Smith was not a big believer in inner alignment.

Higher-level goals do not have to be directly expressed to the components being relied upon to collectively achieve them. In real-world natural and social systems, these goals are often instead connected to the constraints and affordances that actually matter to the components. Depending on the problem itself, on the environment, and on the agents, a goal might be translated into price or physiological signals, into food rewards or grades, into relationships or roles, into laws or traditions.

We can even apply these methods to the things that humans are doing to manage AI. There have been calls to “pace the frontier” in AI development. This could consist of every AI company internally committing to rigorously balancing progress with safety. But as economist Brian Albrecht recently pointed out, we could achieve some of the desired effect by increasing liability. This would naturally shape the incentive structure of AI companies so that their default behavior tends to strike a better balance.

We should rely less on benevolence and on the right people thinking the right things. It’s a lesson Adam Smith tried to teach us, and Friedrich Hayek after him. It’s a tough one for people to learn. But AIs are getting smarter every day, for better or worse. Maybe one day they’ll be able to explain it to us.

Footnotes

  1. For an overview of the characteristic capabilities acquired by cancers, including resistance to cell death, immune evasion, resource acquisition, and invasion/metastasis, see Douglas Hanahan, “Hallmarks of Cancer: New Dimensions” (2022). For more on leader-follower dynamics in cancer, see Kevin J. Cheung and Sally Horne-Badovinac, “Collective migration modes in development, tissue repair and cancer” (2025). ↩

  2. See Friedrich Hayek, “The Use of Knowledge in Society” (1945), on the problem of coordinating knowledge dispersed among many actors. ↩

  3. For an overview, see Michael Levin, “Bioelectric Signaling: Reprogrammable Circuits Underlying Embryogenesis, Regeneration, and Cancer” (2021). ↩

  4. See Francis L. W. Ratnieks and P. Kirk Visscher, “Worker policing in the honeybee” (1989). ↩

  5. See Nikolai Bernstein, Bernstein’s Construction of Movements: The Original Text and Commentaries (2021), edited by Mark L. Latash, on task-organized functional groupings. ↩

  6. See Norbert Wiener, Cybernetics: Or Control and Communication in the Animal and the Machine (2nd ed., 1961). ↩

  7. Adam Smith, The Wealth of Nations, volume 1 (1776). ↩