PixWorld

Unifying 3D Scene Generation and Reconstruction in Pixel Space

arXiv 2607.05373

* Equal contribution    † Corresponding author
1Nanyang Technological University    2AISphere

Abstract PixWorld is a single end-to-end pixel-space diffusion model that unifies 3D scene generation and reconstruction. It partitions posed multi-view inputs into a clean subset (reconstruction) and a noisy subset (generation), processes both with a two-stream diffusion transformer, and decodes a pixel-aligned 3D Gaussian representation in a single forward pass. A pixel-space flow-matching loss is imposed directly on the rendered multi-view images — so optimization is aligned with 3D scene fidelity instead of an intermediate latent target, with no VAE/RAE. A geometry perception loss aligns rendered views with ground truth in the feature space of a frozen 3D foundation model (π³ / VGGT), supplying 3D structural supervision beyond 2D photometric and perceptual losses. After distillation, the 4-step model generates a scene in ~0.6 s — up to ~1000× faster than diffusion-based world generators.

Highlights

~0.6 s
Per scene · 4-step distilled (480P)
~1000×
Peak speedup vs diffusion generators
71.04
WorldScore average
26.21 dB
PSNR · RealEstate10K 4-view recon

Demo Video

Method

3D Reconstruction

Image → 3D Generation

Text → 3D Generation

Comparison with Baselines

Scene Gallery

Indoor

Outdoor

Stylized

3DGS-Rendered Depth (RGB-D)

Inference Speed

Citation

@misc{gao2026pixworld,
    title={PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space},
    author={Sensen Gao and Zhaoqing Wang and Qihang Cao and Dongdong Yu and Changhu Wang and Jia-Wang Bian},
    year={2026},
    eprint={2607.05373},
    archivePrefix={arXiv},
    primaryClass={cs.CV},
    url={https://arxiv.org/abs/2607.05373}
}

Click to copy

×