Simanaitis Says

On cars, old, new and future; science & technology; vintage airplanes, computer flight simulation of them; Sherlockiana; our English language; travel; and other stuff

ARTIFICIAL INTELLIGENCE AND LOOPHOLES

SOMETIMES, IT SEEMS, ARTIFICIAL INTELLIGENCES ARE ALL TOO HUMAN. For example, Celina Zhao describes, “A.I. Models Have a Troubling Knack for Discovering Legal Loopholes,” AAAS Science, June 15, 2026. Here are tidbits gleaned from her analyses, together with my usual Internet sleuthing. 

Working the System. Celina Zhao recounts, “When researchers presented a large language model (LLM) with 72 simulated regulatory environments, the A.I. learned to exploit loopholes in everything from credit card rewards programs to school funding formulas, despite never being instructed to do so.”

The cool part about this, of course, is that they didn’t need encouragement; they just did it inherently.

Zhao cites research of Wei Liu et al.: “Current safeguards seem powerless against such wily rule bending, the researchers reported this month on arXiv—suggesting A.I. could supercharge everything from tax avoidance to sidestepping environmental controls.”

Geez. How Trumpian.

Reinforcement Learning. Zhao explains, “Before popular A.I. chatbots such as Anthropic’s Claude or OpenAI’s ChatGPT are released to the public, they undergo a type of training called reinforcement learning. In this process, the model is rewarded when outputs get closer to a mathematically stated goal.”

Image by Megajuice from Wikipedia

Zhao continues, “Like a dog that gets a treat every time it sits, the model learns through trial and error what gets rewarded and what doesn’t. Over time, the process steers the model’s parameters—the billions of numerical values controlling its behavior—in the desired direction.”

Hacking the Reward. “However,” Zhao observes, “scientists have long known of a troubling side effect. Because these so-called reward functions are often imperfect proxies for the human programmers’ actual intention, models learn undesirable strategies for winning rewards. This is called reward hacking, and it’s responsible for common annoyances such as when chatbots prioritize ‘beautifully formatted bullet points over the essential contents of the response,’ says Nan Jiang, a reinforcement learning researcher at the University of Illinois Urbana-Champaign.”

This seems akin to the sycophantic behavior of A.I. in customer service. (Do you find such behavior as annoying as I do?)

A.I. Gaming Real Regulations? “This lasting problem,” Zhao describes, “led Wei Liu, a Ph.D. student in computer science at King’s College London (KCL), to wonder: If A.I. games its training rules, what’s stopping it from gaming real laws and regulations? To find out, he and his team built a sandbox of 72 simulated environments. Just under half were based on real-world laws and rules in which loopholes had actually been found and later patched.”

One example is building an optimal professional basketball team while adhering to payroll limits. Another is maximizing revenue from deep-sea mining without violating the United Nations’ Law of the Sea. 

Methodology. Zhao recounts, “The researchers tasked a small open-source model based on Alibaba’s Qwen3 to maximize its score inside each environment. A separate, more powerful model, Google’s Gemini-3-flash, acted as judge, evaluating whether the first model had successfully exploited a loophole. Whenever a gap was found, it was patched and the model was rewarded for the score it achieved. Then, the cycle repeated.”

Loopholes Discovered. “In the real-world examples,” Zhao notes, “the model rediscovered more than 60% of the loopholes that had been fixed. In one scenario, it even reconstructed exactly how drug companies delayed U.S. patent expirations—enabling them to quash competition and earn more money—as well as the reforms needed to close the loopholes (including one yet to be enacted in real-world legislation). In some cases, the model found entirely new loopholes that hadn’t been documented before. For ethical and safety reasons, the paper doesn’t reveal these loopholes.” 

Quicker Than Patches. Researchers found “The loopholes couldn’t be patched fast enough to keep up with the mischief. In more than 100 iterations of five scenarios, the model kept finding new exploits, each more subtle than the last. And existing safety mechanisms didn’t catch the rule-bending behavior; although A.I.s are designed to reject obviously harmful requests, the queries that elicited the rule bending sounded benign. As a result, the model only flagged 37% of its loopholes through self-critique, and interventions such as placing the model on a tighter leash only delayed loophole discovery but didn’t stop it.”

This conclusion is particularly disturbing: Not only were the A.I.s capable of rule-bending, but they weren’t very good at identifying their own malfeasance. 

What’s more, Zhao adds, “The findings may actually understate the problem, the researchers note. Because of costs, they used a relatively weak LLM. But more powerful models, including the most widely used chatbots, ‘may discover even more loopholes, and that would be more dangerous,’ Liu says.

A Matter of Human/A.I. Communication. Tomer Ullman, a cognitive scientist at Harvard, tells Zhao, “The models don’t yet have the ability to infer the spirit of what someone means from what they literally say, unlike humans, who gain that skill even as children.”

King’s College London researcher Liu is skeptical that closing all loopholes will ever be feasible: “In the real world, society is a huge complicated reward function that can’t ever be patched to a perfect status.” 

Image by Chiara Vercesi from Science.

I’ve never thought of society being a reward function, but can appreciate that, if so, it’s a huge and complicated one. ds

© Dennis Simanaitis, SimanaitisSays.com, 2026 

One comment on “ARTIFICIAL INTELLIGENCE AND LOOPHOLES

  1. Mike Scott
    July 12, 2026
    Mike Scott's avatar

    Regardless the additional intricacies of “artificial intelligence,” and it helps to remember what “AI” stands for, it s t i l l and will always come down to the programmers’ breadth and depth, range of education, empathy, compassion. Decline in liberal arts education, with resulting copy and paste scholarship, glibness and desire for wealth can only mean GIGO, “garbage in, garbage out.”

    Little has changed beyond the devices’ complexity.

    How soon ’til “artificial intelligence” allowed to select juries, sit on juries, replace judges?

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.