Harnessing Coupled Stream Completion For Human-Object Interaction Modeling
Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories&rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of ea...