Axolotl3D: A Unified Framework for Faithful 3D Shape Completion
Authors: Anita Hu, Maria Shugrina
Submitted: 22 July 2026
Accepted to: ECCV 2026
arXiv: 2607.20660 (cs.CV)
Abstract
Recent 3D generative models leverage large-scale priors and diffusion architectures to produce high-quality geometry from a single image. However, these methods typically assume complete visibility and single-view inputs, which limits their applicability in multi-view, heavily occluded, or editing-based scenarios. While prior works address some of these challenges, they lack a unified framework capable of controllable 3D completion under diverse conditioning signals.
In this paper, we introduce Axolotl3D, a multi-modal, occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud acts as a geometric anchor, promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment within a shared 3D coordinate system. Our unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.
Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings. Additionally, Axolotl3D achieves strong results in real-world reconstruction and geometry-consistent editing, making it a versatile tool for modern 3D vision applications in 2026 and beyond.
via ArXiv CV
