
SH
Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du
· 1 min read
ResearcharXiv cs.AI
ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs
arXiv:2609.32344v1 Announce Type: new
Abstract: For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confidence, and relation metadata; a single ranking supports multiple write budgets while preserving all facts in external memory. On CounterFact with Qwen3-4B, ALLOT reaches 0.760 accuracy at a 20% parametric-write budget and recovers 78.4% of the budget-matched oracle gain, with 80% fewer parametric writes than dual-writing every fact. At this budget, jointly adding retrieval and relation features to text improves normalized oracle gain by 6.2 percentage points. Complementary Qwen3-0.6B shared-store results achieve dual-write-level accuracy with 6-14.5% parametric writes, and cross-benchmark transfer retains approximately 88% of in-domain gain. These results support allocating adaptation capacity according to its incremental value rather than treating every factual update as an equally valuable training target.
Original source
This story was published by arXiv cs.AI and written by Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


