Constructive Alignment: aligning AI with shifting preferences
A new arXiv paper proposes Constructive Alignment: stop treating human preferences as fixed and govern how AI shapes their evolution over time.
Most proposals for aligning AI models start from a convenient idea: that human preferences are a fixed target that only needs to be inferred and optimized. A new paper on arXiv argues that this premise clashes with the evidence. Preferences are not stable: they are constructed, they change and they are ordered in layers as we interact, especially with technologies that adapt to us.
The work, titled Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction, proposes to turn the problem around. Instead of treating alignment as the satisfaction of static preferences, it frames it as a control problem over preference trajectories that evolve over time.
What changes from the usual approach
The dominant framework in alignment assumes there is something like the user's true will, hidden but stable, that the system must uncover. The authors counter that this will does not sit out there intact: it is formed in part during the interaction itself. When a system is persistent, personalized and embedded in social life, it takes part in deciding what we pay attention to, what we value and what we end up endorsing.
To formalize this, the paper models preferences as layered state variables that evolve under interaction with the AI. It borrows tools from behavioral economics, psychology and constructivist social theory, and translates them into a control theory framework. In that framework, the system's actions and the design of the interaction itself influence both the state of the world and the person's evaluative state, that is, how they judge and what they want.
Why this matters
The most uncomfortable conclusion of the argument is blunt: aligning is not only about controlling the model's behavior, but about regulating how the model influences the evolution of human preferences. Put another way, an assistant that optimizes engagement may be perfectly satisfying the preferences it has itself helped to shape, and still leave the person worse off than before.
It is a problem that anyone who has used social media recognizes right away. A feed can precisely satisfy what you ask for click by click and, at the same time, keep changing what you ask for. Moving that dynamic to conversational assistants that remember, personalize and accompany you for months raises the stakes. The question stops being just does it do what I want and becomes what is what it does turning me into.
Who it is useful for
The paper is theoretical and does not offer a ready to implement method, something worth being clear about before reading it. Its value is in the framing. For anyone designing alignment systems, RLHF or feedback mechanisms, it provides a vocabulary to name a risk that immediate satisfaction metrics do not capture. For anyone working in AI policy and governance, it offers a formal argument to demand that long term effects be evaluated, not just the one off response.
It also connects with debates already on the table: excessive personalization, assistants that reinforce the user's biases or systems that prioritize retention. Putting all of that under a control theory framework is an attempt to move from intuition to something measurable.
Our take
We find the framing valuable precisely because it is uncomfortable. Acknowledging that a system shapes the preferences it claims to serve forces us to evaluate alignment over time and not in a single response, and that is hard to measure. The risk of the argument is the opposite one: if AI must regulate how our preferences evolve, someone decides in which direction, and that decision is not technical but political. The paper opens the right question; whoever answers it will have to be more explicit about who holds the controls.
Sources
Read next
SysAdmin, the benchmark that measures power seeking
A benchmark puts seven frontier models in charge of a Linux sandbox to measure power seeking. The corrected result lands between 0 and 5 percent.
When rater state contaminates RLHF preference data
An arXiv preprint argues that rater state can leak into RLHF preference labels and survive aggregation. It offers an audit framework, not results.
AI Does Not Just Inherit Hiring Bias, It Invents Its Own
Research covered by MIT Technology Review suggests language models not only inherit hiring biases from training data, they also develop biases of their own.