Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Authors: Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
Submitted: August 3, 2026 (arXiv:2608.02711, cs.CV)
Abstract
Recent progress in image generation has demonstrated the potential of unified multimodal models capable of integrating understanding, generation, and editing within a single architecture. However, unified 3D modeling remains limited by the scarcity of multimodal data, particularly large-scale, geometrically consistent 3D editing datasets. To address this gap, we introduce Hunyuan3D-Buffalo 1.0, a unified framework that supports 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation—all within a single architecture. To enable scalable training, we construct an 87-million-sample 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while also demonstrating strong understanding and part-generation capabilities. Our analysis further reveals that both generation and understanding improve editing performance, validating the effectiveness of unified 3D multimodal training. As the field moves toward more integrated 3D pipelines in 2026, this work sets a new benchmark for scalable, multimodal 3D content creation.
Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
via ArXiv CV
