Preprint
Aug 2026
Beyond Single Object: Learning 3D Relations with Large Language Models
This work proposes a framework for detailed object-level reasoning across multiple objects with three components: MO3D, an instruction dataset requiring fine-grained multi-object comparison; Multi-3DLLM, using a minimal Patch-Interaction Transformer (PIT) that models inter-/intra-object relationships while preserving local geometry.
K. Ide, Ryousuke Yamada, Yue Qiu et al.
· 0 citations