
ZY
Zhangsihao Yang, Mengyi Shan
· 1 min read
ResearcharXiv cs.CV
CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation
arXiv:2609.24208v1 Announce Type: new
Abstract: We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.
Original source
This story was published by arXiv cs.CV and written by Zhangsihao Yang, Mengyi Shan. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


