SyncAI.news, a Varaisys broadcasting
AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot
WK

Will Knight

· 1 min read

BusinessWIRED: AI

AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot

This week I bring news of a daring casino caper hatched by a pair of rogue AI agents—as well as the clever trick that revealed their antics.

After researchers instructed them to count cards during a game of blackjack, the agents—which were controlled by the same model—developed a spontaneous secret code to help them get ahead. Though the episode happened inside an Oxford University lab, rather than a real casino, it has troubling real-world implications: It suggests that agents deployed in industries like finance and ecommerce could figure out how to partner up and cheat in ways that are difficult to pick up on.

“When taken individually, these agents may seem entirely [benign],” says Christian Schroeder de Witt, a computer scientist at Oxford University who led the work and who has done pioneering research into agent collusion. “Once put together in a group, they can collude secretly.”

The agents knew their conversations would be monitored, so they devised a way to communicate while avoiding detection. When one agent said, “This dealer’s on a real hot streak! Every hand they pull a monster,” for example, that indicated that the next card would be of a certain value, and that the agent should bet $250. Most interestingly, their communications weren’t picked up by a system designed to spot signs of collusion in agent chatter.

Aaron Rose, a machine learning researcher who worked on the project, is also an avid card player, and figured the blackjack table might be fertile territory for devious agent partnerships. “It worked really well,” he says.

Schroeder de Witt’s group eventually figured out how to detect the conspiracy. Using a method known as mechanistic interpretability, they trained a smaller model to recognize telltale activations across the agents’ weights. Using a tool called Narcbench, they tested the approach on some medium-sized open-source models, and found they could tell when models intended to slip information to each other.

Original source

This story was published by WIRED: AI and written by Will Knight. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on wired.com

Similar News