🤝 Open to Collab
Vincent-Daniel Yun
yunuyean
AI & ML interests
None yet
Recent Activity
submitted a paper about 13 hours ago
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs posted an update about 13 hours ago
Can LLMs from different model families directly share KV caches, without receiver-side prefill?
Our answer is HeteroFold.
I am so excited to share our new paper: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs!
Recent work on prefill-free KV cache transfer has shown that LLMs can reuse previously computed context instead of repeatedly performing receiver-side prefill. However, existing approaches have largely focused on models within the same family or with compatible tokenization.
We introduce HeteroFold, a method for prefill-free cross-family KV cache transfer that enables models such as Llama, Qwen, and Ministral to directly share KV caches despite differences in tokenizers, model depth, KV structure, and representation spaces.
HeteroFold aligns tokens across different tokenizers and layers across different architectures, maps the sender’s K/V states into the receiver’s representation space, and calibrates the transferred cache to preserve the receiver’s attention patterns and outputs. Both the sender and receiver remain frozen, and the learned mappings allow the receiver to directly decode from the transferred cache without processing the original context again.
At 32K context length, HeteroFold achieves about 10.7× faster transfer than native receiver prefill. We evaluate all six transfer directions among Llama, Qwen, and Ministral, with strong results across long-context, short-context, and multi-agent settings.
What excites me most is the possibility of making KV caches reusable across model-family boundaries. Instead of different LLMs repeatedly recomputing the same context, heterogeneous models can directly reuse computation from one another.
Paper: https://arxiv.org/pdf/2609.32259 upvoted a paper 25 days ago
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs