MAGICITY4D: Controllable and Editable 4D City Scene Generation Using MLLM-Enhanced Procedural Content Generation
Abstract
In recent years, 3D city scene generation has made significant progress. However, controllable and editable 4D city scene generation remains a largely unsolved challenge. The core difficulty lies in effectively incorporating temporal dynamics. It also requires precise control over scene content and higher system coordination. To address these challenges, we propose MagiCity4D, a multimodal 4D city scene generation framework. By deeply integrating Multimodal Large Language Models (MLLMs) and Procedural Content Generation (PCG), our framework enables both controllable 4D city scene generation and real-time, fine-grained editing without additional training. The experimental results show that MagiCity4D outperforms existing methods in overall performance, demonstrating its potential for advanced city scene generation. Our project page: https://erxucomeon.github.io/MagiCity.

Fig. 1: Pipeline of MagiCity4D. The MagiCity4D pipeline is divided into two stages: Stage 1 (top) manages open-source asset generation and integration, while Stage 2 (bottom) utilizes the outputs from Stage 1 to generate and edit the 4D city scenes.

Fig. 2: Quality comparison.

Fig. 3: Spatial consistency of city Layout.

Fig. 4: Visualization of generated City scenes.

Fig. 5: City scene editing.
Fig. 2 shows the comparison of the generation quality of each method, Fig. 3 shows the visual comparison between the layout sketch (left) and the generated city layout (right), demonstrating our method's high degree of spatial consistency. Fig. 4 shows the visualization of the city scenes generated by MagiCity4D. Additionally, Fig. 5 shows MagiCity4D's powerful ability to edit city scenes.