SyncAI.news, a Varaisys broadcasting
PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory
SY

Seoyoon Yum, Sehoon Kim

· 1 min read

ResearcharXiv cs.LG

PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory

arXiv:2610.04537v1 Announce Type: new Abstract: On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode latency by 60-61% on two M4 systems, whereas prefill latency rises by only 5.7-6.9%. We study LLM phase as an admission signal for independent CPU retrieval under controlled LLM workloads. PHASEGATE calibrates separate concurrency limits for prefill and decode, selecting four and one on our base-M4 configuration. Under a backlogged queue, it achieves 2.0 times the aggregate retrieval throughput of the best tested feasible fixed policy, with both p95 LLM latency metrics within 1.25 times their no-retrieval baselines in all seven held-out runs. A phase-blind control, TimeGate, uses the same two limits on a calibration-derived schedule without observing LLM phase. It achieves similar retrieval throughput but violates the output-token latency limit in every run. M2 and M2 Pro Mac minis reproduce the policy ordering, while output-length sweeps show that the advantage narrows as decode occupies more of each request.

Original source

This story was published by arXiv cs.LG and written by Seoyoon Yum, Sehoon Kim. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News