SyncAI.news, a Varaisys broadcasting
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
VY

Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Sai Praneeth Karimireddy, Murali Annavaram

· 1 min read

ResearcharXiv cs.AI

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

arXiv:2609.32259v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

Original source

This story was published by arXiv cs.AI and written by Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Sai Praneeth Karimireddy, Murali Annavaram. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News