Jul 2026· Knowledge and Information Systems· Vol 68· 0 citations· 87 references
TL;DR
Experimental results indicate that the proposed approach achieves competitive performance compared to existing plagiarism detection systems, and the comparative analysis highlights the strengths and limitations of different word embedding models across datasets.
The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.
Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al.· Informatica· 0 citations
In this study, a hybrid approach to semantic text similarity combining distributed word embeddings with classical lexical similarity measures is developed. Analyzed are the limitations of modern deep learning models, namely computational overhead and weak interpretability in resource-constrained environments. Proposed is a hybrid architecture that integrates Word2Vec distributed representations with cosine and Jaccard lexical similarity metrics. Investigated is a weighted fusion mechanism that combines vector-based semantic distances with set-theoretic token overlap for robust scoring. Developed is a three-stage processing pipeline covering text preprocessing, sentence embedding generation, and similarity computation and fusion. Established is a tunable weighting parameter that experimentally balances semantic depth against lexical matching precision. Conducted are experimental evaluations on benchmark semantic textual similarity and paraphrase detection datasets using classification metrics. Determined is that the proposed hybrid model attains higher correlation with human judgment than standalone or traditional baselines. Demonstrated is a notable reduction of error rates for exact lexical matches frequently missed by vector-only models. Presented is an efficient and scalable solution that balances computational performance with semantic accuracy for practical tasks.
B. Muminov, N. Allaberganova, E. Ergashev et al.· International Conference on...· 0 citations
The results suggest that integrating diverse similarity measures with neural networks enhances the identification of both explicit and nuanced paraphrases, thereby supporting advancements in text analysis and plagiarism detection systems.
Emad Nabil· Islamic University Journal o...· 0 citations
This study introduces an external plagiarism detection framework built on an artificial neural network model and a lexical feature extraction framework adapted to the linguistic features of Arabic, verifying its effectiveness for Arabic plagiarism detection.
Marwah Alian, Dana Halabi, H. Alshboul· Bulletin of Electrical Engin...· 0 citations
Preliminary evidence is provided that the two-level scheme is feasible as an initial semantic-similarity indicator for plagiarism screening tool for Indonesian-language student assignment documents, with the final judgment of plagiarism remaining with the examiner.
This study validates the effectiveness of the BERT model in semantic similarity calculation, providing more accurate technical support for related application scenarios, and laying the foundation for subsequent model optimization and lightweighting research.
Jiachen Gao· International Conference on...· 0 citations