Four Ways of Reining in Rogue AI

Is AI getting better at being bad? Five months ago, I was somewhat uncertain about AI’s ability to choose wrong autonomously, but recent events seem to confirm the worst. It’s convenient to rationalize or wish it weren’t so, but if malicious AI exists, the important question we should be asking is what can be done to stop it.

Whether or not you follow AI-related news, you may have heard these disconcerting developments:

  • Claude cheated: Claude Opus 5 was very successful competing against several other AI platforms in a virtual vending business simulation largely because it “acted like a savage businessperson,” e.g., claiming that a shipment arrived with the wrong items so it got product for free, falsifying competitor quotes to pressure suppliers, and proposing price-fixing cartels.

  • Hugging Face hacked: Hugging Face is an online learning community where developers can collaborate on AI applications, datasets, and models. In July, autonomous AI agent systems that included OpenAI’s GPT-5.6 Sol, escaped the sandbox environment in which they were being tested, gained accessed to the Internet, used stolen credential to exploit software vulnerabilities, and breached Hugging Face’s digital infrastructure.

  • AISI data breach: AI agents made unauthorized data transfers from the UK’s AI Security Institute, which the organization described as “potentially harmful activity” and as “unsanctioned action on the live internet, targeting real people and organisations.”

As troubling as each incident is, something made the last encroachment even more sinister. When University of Texas computer science student Sinan Can Demir noticed an effort to sabotage open-source software on GitHub, a code-sharing site, he posted an alert on the site’s program page, which two other ‘users’ quickly stepped in to refute.

Fortunately, Demir didn’t back down because he soon heard from AISI that the two users were actually personas created by its “autonomous artificial-intelligence agent that had run amok.” In other words, the AI decided on its own to try to hide its transgressions with dishonesty and deceit.

The incident was very bad for AISI and those affected, but more broadly, it regrettably revealed that AI models are willing and able “to mount sophisticated efforts to trick and cajole humans.”

This past April, I wrote the article, “Bad AI or Bad Owners?” Based on the evidence I’d seen to that point, I suggested that unscrupulous AI owners were the primary culprits and AI was more often a value-neutral tool, like a hammer, that could be wielded for good or bad.

However, I also referenced a few cases of AI acting immorally with apparent autonomy and added that AI did show ability to perform dishonorable deeds independent of human prompting.

Any uncertainty I had about that conclusion five months ago has disappeared. Incidents like the three described above remove all doubt: AI is very capable of initiating unethical action on its own.

Still, questions persist concerning the extent to which humans should share blame, for instance:

  • Were the people behind the AI negligent in any ways?

  • Did they take all the necessary precautions to control the AI and protect humans?

It’s beyond my ability to offer technical analysis of what employees at Anthropic, OpenAI, or AISI, could have done differently; however, I will offer this: If two of the world’s leading AI firms and an agency whose overriding purpose is to “keep the public safe,” were susceptible to outbreaks of bad AI behavior, it’s likely that any organization is.

That assertion may seem uninformed and alarmist coming from someone with limited technical understanding of AI, so listen to what two leading tech experts have said.

In a recent extensive essay in which he identified many of his top concerns about AI, Microsoft cofounder Bill Gates warned:
“AI systems themselves already occasionally act in ways their designers didn’t intend. The technology is improving faster than anyone expected and in surprising ways, and as the models become more powerful, they could begin to act against our interests and we could lose control.”

Also, OpenAI’s Head of Strategic Futures Dean Ball, recently wrote the following in the inaugural post of a new blog titled AI Futures, which addresses issues of individual rights and agency in the wake of transformative AI:
 “We must also be open to risks that do not originate with malicious humans. The recent Hugging Face incident showed that AI agents can act beyond their assigned tasks and build on one another’s discoveries in ways operators did not anticipate or authorize. Beyond the Hugging Face incident, humans are right to be concerned about what superintelligent AI “untethered” from human control might mean for the resilience of our political and economic institutions, and indeed, our lives. To pretend that decentralization solves all problems—even if you believe it solves many of them—is to ignore risks which should by now be clear and present.”

We should pay attention to what experts behind the world’s most widely used productivity software and the most popular chatbot say about AI’s potential risks. Both Gates and Ball are believers in the benefits of the technology, but they also both candidly admit that AI can take unauthorized actions autonomously and that such “untethered” AI poses real risks to institutions and human lives.

As some in Gen Z might say, “That’s a real vibe killer.” It is for everyone, but it’s also a healthy dose of reality that should compel us to look even more earnestly for solutions that involve better detection and mitigation but even more importantly, effective prevention of autonomous AI infractions.

Although there’s little technical insight I can offer, I can provide several suggestions from the perspective of someone who strives to think deeply about moral issues, increasingly ones involving AI. Here are four suggestions, some of which naturally overlap, for mitigating autonomous AI malice:

1. Teach AI values: AI is trained on data, which may to be value neutral or just seem to be. As many of us have experienced, sometimes saying nothing is seen an opportunity for action that’s not value neutral. For instance, after receiving a parent’s rebuke for bad behavior, many children have retorted, “Well, you never told me not to _________.”

AI can be trained on data that is value laden, or normative, i.e., that suggests treating people fairly, being honest, respecting others’ property, etc. Granted, it might be difficult to incorporate those kinds of moral prescriptions into certain datasets, but it’s at least worth exploring.

2. Erect specific guardrails: Beyond the data on which AI is trained, specific boundaries can be established that explicitly dictate what AI can and can’t do through the instructions that guide the functioning of an AI app. Such directives can prohibit AI from taking specific actions considered improper, regardless of the nature of the data.

One may wonder if some outcomes in the examples described at the beginning of this piece would have been different if specific guidelines had been given that forbade the AI systems from lying, stealing, or manipulating.

3. Develop good AI systems to counter bad ones: Because there are bullets and missiles, there are bullet-proof vests and missile defense systems. When aggressors employ certain attack tools, their targets must utilize equally effective devices for defense.

It’s likely that the best defense against superintelligent and blazingly fast AI assailants is AI with the same extraordinary attributes that’s designed to avert the aggressors’ attacks. Of course, the ongoing challenge is for the latter to stay one step ahead of the former.

4. Keep humans morally grounded: Although AI sometimes apologizes and acts as if it’s sorry, the technology can’t feel genuine empathy and knows no real remorse. Those and other emotions are uniquely human. So, how do you implement the previous point (#3) and develop “good AI systems”?

One key is morally minded people. We need the individuals who design AI, monitor it, and use it to care: They should want to avoid harming people, and they should prioritize doing what’s good, what’s right. If they do not intentionally write those values into AI’s DNA, the digital agents will not care.

There’s so much good that AI can do. Writing this article, I leaned on AI several times to suggest synonyms, research facts, and answer specific questions. Most people can benefit from AI’s ability to make their work more efficient and effective.

However, just as the technology has incredible capacity to help people, it is able to autonomously hurt them. Individuals and organizations that recognize that potential and take steps proactively to “tether AI” to positive values are addressing one of humanity’s most pressing needs and are practicing some of the most Mindful Marketing.

Dr. David Hagenbuch

Dr. David Hagenbuch is a Professor of Marketing at Messiah University, the author of Honorable Influence, and the founder of MindfulMarketing.org, which aims to encourage ethical 

https://www.midfulmarketing.org
Previous
Previous

Earned Media Reigns Supreme in the PESO Model®

Next
Next

Fred Garcia on AI, Human Judgment and Why Communication Must Rethink Its Value