
PV
Paraskevas V. Lekeas, Giorgos Stamatopoulos
· 1 min read
ResearcharXiv cs.AI
How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models
arXiv:2604.27167v3 Announce Type: replace-cross
Abstract: On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, lock at full cooperation, the metric's maximum distance from Nash with zero variance across replicates, while Llama-3-8B plays near-Nash. Opening the models, a logit-lens analysis finds a distributed cooperative override. Intermediate readouts lean toward the Nash action through roughly three quarters of network depth before a late surge toward cooperation, and the final layer settles the contest. The size of that final correction, not the surge, rank-matches chain-of-thought behavior across scale and two architectures. In the 8B the override is a single causally controllable direction in the residual stream; steering it dials the decision, and clamping its component at one position of one layer moves the choice strictly monotonically, Spearman rho = 1.000, with generation fluent. The circuit is lexical. It survives name removal and payoff rescaling but disengages when Cooperate and Defect are replaced with neutral labels, and on 48 payoff-random games with neutral surfaces no model locks cooperative on any dilemma or shows general equilibrium competence. In mixed-model populations a single Nash-playing agent collapses cooperation contagiously. What suppresses Nash play in large language models is a word-triggered circuit rather than missing competence, and it can be measured, bounded, and controlled.
Original source
This story was published by arXiv cs.AI and written by Paraskevas V. Lekeas, Giorgos Stamatopoulos. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


