
CM
Chika Maduabuchi
· 1 min read
ResearcharXiv cs.CV
Panoptic Scene Program Diffusion Transformer
arXiv:2609.31780v1 Announce Type: new
Abstract: Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.
Original source
This story was published by arXiv cs.CV and written by Chika Maduabuchi. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


