This work defines a modular late-fusion function that combines semantic similarity (cosine of embeddings) and structural similarity (bibliographic coupling) with a tunable weight alpha whose value is chosen according to the specific task.
Abstract
Structural graph analysis of the academic publishing network captures the topological relationships between entities but does not see the content of works. Building on our structural approach, this work complements it with a semantic layer and a parameterized structural-semantic fusion. We represent scientific documents by citation-informed vector embeddings (SPECTER2) and store them in an embedded vector database keyed by the stable OpenAlex ID, so that they connect directly to the graph layer. We define a modular late-fusion function that combines semantic similarity (cosine of embeddings) and structural similarity (bibliographic coupling) with a tunable weight alpha whose value is chosen according to the specific task. On the corpus of VSB - Technical University of Ostrava we show two things: citation-informed embeddings agree with the expert OpenAlex topical taxonomy better than a TF-IDF baseline, and in a recommendation use case the structural, semantic, and combined signals carry information in different regimes depending on the available data. Hybrid fusion here is not a universally better method but an explicit mechanism for steering complementary signals according to the task. We release the whole approach as an open-source extension of the apnet library with a reproducible workflow.
The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Buil...
Citation recommendation plays a critical role in scholarly information retrieval by assisting researchers in identifying relevant and influential literature. Existing approaches typically rely on either textual semantic modeling or graph-based citation analysis, but often fail to jointly capture semantic relevance, str...
Jia-Bo Liu, Hua-Xiong Zhang· Journal of universal compute...· 0 citations
This paper revisits semantic projections and related count-based representations as interpretable directional semantic structures for semantic analysis in document corpora and web-based information environments and demonstrates that semantic projections effectively capture persistent contextual structures while remaini...
Mabel López-Bordao, Antonia Ferrer-Sapena, Pablo Lara-Navarra et al.· Information· 0 citations
Systematic literature reviews (SLRs) face challenges from rapid publication growth and low-quality AI-generated content. Simple database queries often retrieve publications that are not thematically coherent, making meaningful clustering difficult. This study aims to develop and evaluate a hybrid method to automate clu...
Sebastian Matysik, Joanna Wiśniewska, Paweł Karol Frankowski· IEEE Access· 0 citations
Scientific papers may relate by problem, method, result, or contribution, but document-level retrievers collapse these into a single similarity score without saying why they are related. Citation- and similarity-based retrieval alone also confines search to the neighbourhood of what is already known, whereas generative...
Italo Luis da Silva, Hanqi Yan, Yujing Wang et al.· 0 citations
These results demonstrate that collaborative modeling of instance-level and batch-level structural context can effectively enhance structure-aware entity representation and improve fine-grained entity prediction.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.