
AR
Adrien Ramanana Rahary, Nicolas Dufour, Patrick P\'erez, David Picard
· 1 min read
ResearcharXiv cs.CV
Diffusable Latents from Structure-Agnostic Distillation
arXiv:2609.39657v1 Announce Type: new
Abstract: Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at https://github.com/AdrienRR/structure-agnostic-distillation and https://kyutai.org/blog/2026-09-28-structure-agnostic-distillation/.
Original source
This story was published by arXiv cs.CV and written by Adrien Ramanana Rahary, Nicolas Dufour, Patrick P\'erez, David Picard. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


