SyncAI.news, a Varaisys broadcasting
Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
YM

Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao

· 1 min read

ResearcharXiv cs.LG

Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation

arXiv:2609.35333v1 Announce Type: new Abstract: Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.

Original source

This story was published by arXiv cs.LG and written by Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News