SyncAI.news, a Varaisys broadcasting
AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
YS

Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su

· 1 min read

ResearcharXiv cs.CL

AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

arXiv:2606.02461v3 Announce Type: replace-cross Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate experience across a stream of tasks, improve over time, and avoid interference from irrelevant experiences. Unfortunately, existing benchmarks struggle to evaluate this problem rigorously. Most efforts focus on retrieval and reasoning over long-context conversations or documents, while recent lifelong-adaptation benchmarks often stream existing datasets whose cross-task dependencies are unknown, making it difficult to understand what an agent learns and reuses over time. This paper presents an evaluation framework AgentCL for continual learning in agents, centered on distinguishable task streams and metrics for CL properties. AgentCL distinguishes three cross-task relations by the availability of reusable knowledge. In dependent streams, later tasks can reuse knowledge from earlier ones. In conventional streams, whether such reuse exists is unknown. Besides, independent tasks are held out with little transferable knowledge. We use the benchmark to evaluate non-parametric memory designs for continual learning. To diagnose how memory design choices affect continual learning, we develop MemProbe, a probing method that stores interactions, insights, and skills, while filtering unreliable experiences during consolidation. Empirical analysis across coding, deep research, and language understanding/reasoning tasks shows that conventional streams offer limited ability to distinguish memory designs, whereas dependent streams more clearly distinguish their plasticity. Meanwhile, conventional streams and independent tasks often yield limited gains and can expose memory-induced degradation. These results highlight the need for stronger memory designs that balance plasticity and stable reuse.

Original source

This story was published by arXiv cs.CL and written by Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News