Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this...
Si-Ran Peng, Tian-Shuo Zhang, Tian-Yu Fu et al.· 0 citations
Inferring apparent personality from facial images is important in social scenarios for embodied agents in human-robot interaction. Unlike inferring intrinsic personality traits via conversation, this task models first-impression personality perception based solely on facial appearance before interaction begins. Existin...
Self-Supervised Skill Optimization (SSO) is introduced, a comparative framework that learns a reusable skill from unlabeled task instances alone and outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks.
Siran Peng, Cui-Yu Yang, Tianyu Fu et al.· arXiv.org· 0 citations
Private Face Distillation is proposed, an identity-decoupling and geometry-preserving framework that uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relat...
Shuhuan Chen, Xiangyu Zhu, Weisong Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.