
XL
Xiaoze Liu, Ruowang Zhang, Amir H. Abdi, Michel Galley, Zhikai Chen, Siheng Xiong, Xiaoqian Wang, Jing Gao
· 1 min read
ResearcharXiv cs.CL
Do Proactive Agents Need an LLM to Decide When to Act?
arXiv:2605.30152v2 Announce Type: replace
Abstract: Proactive assistants continuously decide when to intervene and what context should support the intervention. Large language model (LLM) pipelines repeatedly interpret activity histories to make these decisions, paying an inference cost even when the assistant remains silent. We show that a small graph model can handle both decisions and improve the language agents it controls. Our key insight is that user activity has a native graph structure: events involve persistent entities whose recurrence connects interactions over time. Triggering and context selection map directly to predictions on event and entity nodes. We introduce a temporal-graph-learning (TGL) controller that learns these predictions jointly and supplies both outputs in one forward pass. The downstream language agent generates suggestions on triggered events using the activity history and scored entities. At 11.13 ms per event on a GPU server, TGL achieves the highest AUCs among nine trigger architectures and gives approximately $4$--$7\times$ trigger-stage speedups over the two single-forward LLM triggers. A shared TGL model improves F1 across all 14 downstream backbones by a mean of 16.7 points. The controller also runs at 13.99 ms on a consumer laptop with an approximately 220 MiB BF16 resident footprint, bringing effective proactive control to on-device deployment.
Original source
This story was published by arXiv cs.CL and written by Xiaoze Liu, Ruowang Zhang, Amir H. Abdi, Michel Galley, Zhikai Chen, Siheng Xiong, Xiaoqian Wang, Jing Gao. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


