Axolotl3D: A Unified Framework for Faithful 3D Shape Completion

3d shape completionaxolotl3deccv 2026multi-modal generationocclusion-awarepoint cloud conditioning

Axolotl3D: A Unified Framework for Faithful 3D Shape Completion


Authors: Anita Hu, Maria Shugrina

Submitted: 22 July 2026

Accepted to: ECCV 2026

arXiv: 2607.20660 (cs.CV)


Abstract


Recent 3D generative models leverage large-scale priors and diffusion architectures to produce high-quality geometry from a single image. However, these methods typically assume complete visibility and single-view inputs, which limits their applicability in multi-view, heavily occluded, or editing-based scenarios. While prior works address some of these challenges, they lack a unified framework capable of controllable 3D completion under diverse conditioning signals.


In this paper, we introduce Axolotl3D, a multi-modal, occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud acts as a geometric anchor, promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment within a shared 3D coordinate system. Our unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.


Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings. Additionally, Axolotl3D achieves strong results in real-world reconstruction and geometry-consistent editing, making it a versatile tool for modern 3D vision applications in 2026 and beyond.

via ArXiv CV

Related