SyncAI.news, a Varaisys broadcasting
Embedding Prediction Helps Image Generation
SX

Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

· 1 min read

ResearcharXiv cs.CV

Embedding Prediction Helps Image Generation

arXiv:2610.02203v1 Announce Type: new Abstract: In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.

Original source

This story was published by arXiv cs.CV and written by Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News