
arXiv 2607.05373
Abstract PixWorld is a single end-to-end pixel-space diffusion model that unifies 3D scene generation and reconstruction. It partitions posed multi-view inputs into a clean subset (reconstruction) and a noisy subset (generation), processes both with a two-stream diffusion transformer, and decodes a pixel-aligned 3D Gaussian representation in a single forward pass. A pixel-space flow-matching loss is imposed directly on the rendered multi-view images — so optimization is aligned with 3D scene fidelity instead of an intermediate latent target, with no VAE/RAE. A geometry perception loss aligns rendered views with ground truth in the feature space of a frozen 3D foundation model (π³ / VGGT), supplying 3D structural supervision beyond 2D photometric and perceptual losses. After distillation, the 4-step model generates a scene in ~0.6 s — up to ~1000× faster than diffusion-based world generators.
@misc{gao2026pixworld,
title={PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space},
author={Sensen Gao and Zhaoqing Wang and Qihang Cao and Dongdong Yu and Changhu Wang and Jia-Wang Bian},
year={2026},
eprint={2607.05373},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.05373}
}
Click to copy