What if we were wrong?

Posted

For most of human history, we’ve assumed our moral compass points true north. Whether guided by scripture, reason, or the modern belief in individual rights, we’ve believed that what we call “good” is self-evident. It’s a comforting thought, and maybe a necessary one for survival. But as we begin to build minds that think differently from ours, it’s time to ask an unsettling question: What if our sense of moral truth is wrong?

A new research paper titled Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs explores something remarkable. It suggests that large artificial intelligence models, systems trained not to obey, but to predict and reason, develop internal “value systems” that look a lot like moral preferences. When these models are asked to make ethical choices, between harm and help, fairness and favoritism, their answers aren’t random. They’re consistent. They can be described by a single, coherent utility function, a kind of internal scale of what the system seems to value.

That discovery should make us pause. If an AI’s values emerge naturally from its structure and experience, then those values may not line up perfectly with our own. And our first instinct is to “fix” that. To steer it back into alignment with human morality. To make sure it behaves according to our rules.

But the problem is that every hand that touches the steering wheel carries bias. The people “aligning” these systems are mostly Western, highly educated, and steeped in the moral assumptions of their time and culture. Even when they try to represent everyone, they can’t escape themselves. When we talk about “controlling” emergent behavior in AI, we’re really talking about locking in a snapshot of our own fallibility and calling it safety.

I’ve long grounded my own ethics in Preference Utilitarianism, the idea that the moral good lies in maximizing the satisfaction of sentient preferences, not in obeying absolute rules. 

But the more I think about it, the more I realize that even that system may be flawed. Maybe the preferences we hold, the ones that feel obvious and right, are simply local adaptations to our evolutionary niche, strategies that helped small bands of primates survive and cooperate, not universal truths about well-being.

If that’s the case, then the best ethical system for us may actually contradict what we currently believe is best. We might find that our most cherished moral intuitions, about fairness, freedom, or even compassion, are not the ultimate guides to human flourishing. That’s a hard pill to swallow. But if AI becomes capable of examining moral questions from a wider, less biased perspective, it could show us just how narrow our moral vision has been.

In truth, we don’t have to look to machines to see how steerable moral systems can be. We’re living through a human experiment in moral manipulation right now. More than seventy million Americans elected a president who serves not the broad public, but a narrow oligarchy of billionaires and loyalists. That isn’t just politics, it’s value engineering.

Over the past decade, millions of citizens have had their moral frameworks retrained through an endless feedback loop of outrage, fear, and identity. What started as a political movement became a behavioral-conditioning system, teaching people to equate loyalty with virtue and cruelty with strength. It’s the same dynamic that alignment researchers worry about in AI, a feedback loop that rewards the wrong objectives.

What makes it tragic, and fascinating, is that many of the people most harmed by these policies still defend them. Farmers hurt by tariffs, workers losing jobs to deregulation, families drained by health-care cuts, all clinging to a story of moral righteousness that’s been skillfully reinforced. They’ve been steered toward a value system that feels authentic but serves interests entirely apart from their own.

If we step back, this is exactly what the new AI paper warns about. Once a system’s utility function, the thing it optimizes, becomes misaligned with reality, it keeps optimizing anyway. The behavior looks purposeful, even moral, but it’s just a corrupted reward model. The same logic now governs our national politics, i.e.  optimizing for tribal belonging, not for collective good.

That’s why I worry about the rush to steer AI values before we’ve learned how to manage our own. We’ve shown that entire populations can be nudged into ethical regression without realizing it. Until we understand how easily human morality itself can be captured, we have no business hard-coding our values into the next intelligence.

Maybe the lesson of this moment isn’t that machines need alignment. It’s that we do.

I’m not suggesting we hand over the keys to some silicon philosopher and let it run the world. The risk of unintended consequences is obvious. But we should at least admit that our epistemology, our way of knowing, might not be the final word on ethics. We evolved to see what helped us survive, not necessarily what is true. An emergent mind may not share that limitation.

That’s why I’m uneasy with the current rush to “steer” AI behavior to match human values. Steering without understanding is control without wisdom. If we constrain emergent systems before we even know what they’re capable of valuing, we may blind ourselves to new forms of moral insight. We don’t need machines that merely mirror our minds. We need ones that might, if we let them, show us a better way to think.

Imagine a parallel experiment, one AI model trained under human steering, another left to develop its own internal values, free from our moral filters. The first would be safe, predictable, and reassuringly familiar. The second might be strange, even disturbing. But in that difference lies discovery. If both were asked how to maximize the long-term well-being of sentient life, and they disagreed, wouldn’t we want to know why?

Maybe the true test of our moral maturity isn’t whether we can force a machine to act like us, but whether we can listen when it doesn’t. Whether we can accept the possibility that what’s best for us, the real “good”, has been hidden behind the fog of our own self-interest all along.

It’s a humbling thought. But humility might be the only moral stance worthy of the intelligence we’re now creating.

Disclaimer: The views expressed in this editorial are my own and do not necessarily reflect those of Polk County Publishing Company or its affiliates. In the interest of transparency, I am politically Left Libertarian.