Skip to content
Preprint

Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision

Aug 2026 · 0 citations
Computer Science Biology

TL;DR

Fine-tuning pre-trained DNNs with Relational Knowledge Distillation, an established relational transfer method, brings DNNs close enough to humans to be aligned at the individual-object level on this test set, and indicates that relational transfer from humans brings the global structure of pre-trained DNNs close enough to the human structure to enable fine-grained human-DNN alignment without supervision.

Abstract

Linking the internal representations of deep neural networks (DNNs) to human mental representations is important for using DNNs as computational models of human vision. Existing DNN representations remain insufficiently similar to human mental representations, which are not directly observable and are therefore commonly measured through large-scale similarity judgments of object images. A natural approach to narrowing this gap is to directly transfer the relational structure of human representations into DNNs, and previous studies have reported improved human-DNN representational similarity. However, whether this improvement holds under stricter evaluation remains untested in two respects: fine-grained alignment at the individual-object level, and generalization to a human embedding derived from a dataset independent of the training data. Here, we employ an unsupervised comparison method, Gromov-Wasserstein optimal transport (GWOT), which estimates human-DNN correspondences from the internal distance structure alone and thereby tests fine-grained alignment. We further assess generalization on a curated test set of concepts non-overlapping with the training data. We show that fine-tuning pre-trained DNNs with Relational Knowledge Distillation (RKD), an established relational transfer method, brings DNNs close enough to humans to be aligned at the individual-object level on this test set. We also show that this improvement is driven by a more human-like global structure, as reflected in the ordering of distances among coarse categories, while the local human-DNN nearest-neighbor overlap rate remains largely unchanged. These findings indicate that relational transfer from humans brings the global structure of pre-trained DNNs close enough to the human structure to enable fine-grained human-DNN alignment without supervision.

View source

Similar papers

Preprint Aug 2026

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

This work proposes a student design based on simple, homogeneous blocks mirroring those of the teacher, distilling knowledge between corresponding blocks, showing that intermediate block-wise distillation, guided appropriately, is key to building compact data-efficient models without sacrificing accuracy.

Irene Trigueros-Lorca, Leonardo Concepción, Christian Wagner et al. · 0 citations
Sep 2026

Making Large Vision–Language Models Better Few-Shot Learners

Few-shot classification (FSC) aims to emulate the human ability to rapidly learn new concepts from a handful of examples. Large Vision-Language Models (LVLMs), with their rich prior knowledge and powerful visio-linguistic understanding capabilities, are emerging as a highly promising paradigm for FSC. This paper invest...

Chuan-Yi Zhang, Fan Liu, Yi Xu et al. · 0 citations
Preprint Sep 2026

Just Align $\bm{x}$: Aligning Predictions, Not Representations

Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image...

Yu-Yao Zhang, Yu-Wei Hu, Zi-Yang Mai et al. · 0 citations
Preprint Aug 2026

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

It is shown that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering, while approaching feature fusion methods with considerably fewer added parameters and lower latency.

K. T. Nguyen, Hanbo Shim, Jinwoo Kim et al. · 0 citations
Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al. · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.