SciRep: A Ranking-Aware Representation Model for Scientific Text
Existing scientific text representation methods based on contrastive learning typically adopt a binary classification paradigm of positive and negative samples, which struggles to capture the complex, hierarchical semantic similarity relationships inherent in scientific texts. To address this, we propose SciRep, a novel two-stage ranking distillation framework. In the first stage, we distill knowledge from a large language model to a medium-scale representation model using generated ranking samples; in the second stage, a multi-teacher strategy further transfers fine-grained ranking capability to a lightweight model. Evaluated on a scientific literature semantic embedding benchmark comprising three tasks, SciRep outperforms the strongest baseline by 11.3% in terms of Average Rank and also achieves the highest Mean Reciprocal Rank scores across all three tasks. These results demonstrate that the proposed ranking-aware distillation mechanism significantly enhances scientific text representation quality while maintaining efficient inference, offering a more effective contrastive learning method for domain-specific retrieval tasks.