Skip to content
Open access

Ontology-Based Text Summarization for Marathi Legal Documents

Jul 2026 · Indian Journal of Science and Technology · Vol 19, pp. 1639-1647 · 0 citations

TL;DR

This research aims to create an ontology-based method for summarizing government documents by improving understanding using domain-specific terms and their relationships and shows that the summaries are accurate, meaningful and match the reference summaries.

Abstract

Objectives/Background: Government documents of Maharashtra state are written in Marathi language, and they are long in length, detailed, and difficult to read quickly. There is a requirement to develop a system which will automatically summarize these documents to understand the key information quickly. Therefore, this research aims to create an ontology-based method for summarizing government documents by improving understanding using domain-specific terms and their relationships. Method: This study proposes a novel technique that uses ontology as a knowledge base, which consists of significant concepts associated with a specific subject and their interconnections. It helps the system to understand the content with its meaning. With this understanding, the algorithm selects important information and generates summaries that maintain the clarity and importance of the text. Findings: The proposed method achieves Precision of 0.82, ROUGE-1 score of 0.85, and ROUGE-2 score of 0.78. These results show that the summaries are accurate, meaningful and match the reference summaries. The proposed system may also reduce the effort required to read long documents by generating short and meaningful summaries. It improves both the quality and reliability of text summarization and works well for domain-specific applications where meaningful summaries are important. Novelty: This research work is different from the traditional statistical text summarization techniques as it uses the ontology to capture text relationships resulting in an improved context understanding and accurate summaries. Keywords: Text Summarization, Ontology-based, Domain knowledge, Marathi Language, Administrative documents

Read PDF

Similar papers

Open access Jul 2026

Ontology-Based Semantic Normalization of Resumes for Classification

Rather than scaling performance uniformly across the entire evaluation suite, the ontology layer acts as a targeted traceability and semantic refinement filter that contributes information beyond filtered-profile selection alone and produces a metric-dependent change in classifier behaviour at the validation-selected threshold.

V. Anghel, Theodor Borangiu, S. Raileanu et al. · 0 citations
Jul 2026

Efficient Scientific Paper Summarization Using Unsupervised Extraction and Transformer-Based Abstraction

The growing volume of scientific literature has driven the demand for automated text summarization systems that are natural and factual. Extractive Text Summarization methods are factually accurate in meaning; still, they can lead to a summary that is not cohesive. On the other hand, abstractive summarization systems improve readability but may introduce factual bias. The paper overcomes these shortcomings by creating a hybrid text summarization system that combines extractive and abstractive methods to maximize both quality and factual content. The framework uses two unsupervised extractive models, HipoRank and PacSum, to extract important sentences, which are then synthesized with the original input document's introduction section and subjected to long-document transformer models, PEGASUSX and LED, to generate abstract-style summaries. Among the tested combinations, the HipoRank-LED configuration achieved the most balanced performance, with ROUGE-1: 0.440, ROUGE-2: 0.220, and ROUGE-L: 0.410 on the PubMed dataset. This combination occasionally produced summaries with greater abstractiveness than the human-written references. Various experiments across the ScisummNet, ArXiv, and PubMed datasets show that hybrid configurations are always better than extractive and abstractive ones. HipoRank-LED is the most efficient model, with ROUGE-1 = 0.440, ROUGE-2 = 0.220, and ROUGE-L = 0.410 on PubMed. Results indicate that combining extractive grounding with long-context transformers improves informativeness and coherence and reduces hallucination errors. The introduction-guided structured input also provides better global context for summarizing complex scientific documents. The findings indicate that the transformer-based abstraction, combined with an extractive text summarization approach, can be a very useful, scalable, and domain-independent model for approximating long scientific texts.

Grishma Sharma, Aditi Paretkar, Deepak Sharma · 0 citations
Preprint Jul 2026

Annotating Topical Legal Insights from Case Proceedings

In this paper, we mainly concentrate on finding concepts or topics from the legal case proceedings, since adopting a structured representation for legal documents, as opposed to a mere bag-of-words flat text representation, can significantly enhance processing capabilities. To achieve this objective, we put forward a set of diverse concepts for legal case proceedings. With this motivation, we propose LeDA, a system for Legal Data Annotation. The system offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface. A novel feature of our system is that it allows to dynamic create new tags for annotation, which is a particularly useful provision for situations where there exists no pre-defined ontology for the entities (concepts) that need to be annotated - these being rather discovered by annotators as they continue examining more documents. The system that we demonstrate is currently in use to annotate a set of concepts from legal documents to construct semantic representations of documents as bags of concepts that can then be used for several downstream tasks, such as prior case retrieval, judgment prediction, and so on. Along with the system features in general, we also describe how LeDA was used by 3 assessors to annotate and adjudicate legal concept names from Indian Supreme Court case proceedings.

Subinay Adhikary, Dwaipayan Roy, Debasis Ganguly et al. · 0 citations
Preprint Aug 2026

ITL: Interpretable Document Alignment with Structured Reference Frameworks

Measuring alignment between documents and structured reference frameworks requires identifying conceptual evidence distributed throughout the text and reporting it through measures that are quantitative, interpretable, and traceable. Many commonly used retrieval and classification approaches return either pairwise similarity scores or one or more class labels, whereas fewer methods provide concept-level scores that are directly traceable to the terminological evidence supporting them. We present \emph{Intelligent Target Locator} (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a \emph{Structured Reference Document} ($SRD$). From the $SRD$, ITL induces concept-specific terminological profiles built from independent terms, bigrams, trigrams, and co-occurrences. Each term is assigned an importance weight that combines concept membership, term-type specificity and inter-concept discriminability. The output is a textual-unit--concept affinity matrix that can be aggregated at different levels of granularity. We conduct an internal consistency assessment using the 17 Sustainable Development Goals (SDGs), evaluating each official goal statement against the $SRD$ induced from the same set of descriptors. Every statement reached its highest affinity with the corresponding concept, and the mean affinity across the remaining concepts stayed marginal relative to the mean reference affinity. This separation indicates that ITL distinguishes the conceptual profiles of the framework. ITL thus offers a general basis for quantifying document alignment with structured frameworks while keeping each result traceable to the terminological evidence that supports it.

R. Giráldez, Dayrelis Mena, Jesús S. Aguilar-Ruiz · 0 citations
Review Open access Jul 2026

Resolution of Expression of Concern: Ontology Features-Based Arabic Text Augmentation Using Word2Vec

Editor's Note: This Resolution notice concludes the investigation initiated by the Expression of Concern (DOI: https://doi.org/10.52866/2788-7421.1386) regarding the Original Article (DOI: https://doi.org/10.52866/2788-7421.1302). The publication timeline is as follows: Original Article → Expression of Concern → Resolution. Resolution of Expression of Concern: The editorial team of the Iraqi Journal for Computer Science and Mathematics has completed a thorough post-publication review of the above-mentioned article, the supporting data, and the authors’ responses to the queries raised. Upon careful assessment, we have determined that the concerns have been addressed and clarified to the satisfaction of the editors. We have found no evidence of misconduct or invalidity in the research. The gindings presented in the original article are considered robust and accurate. This notice formally resolves the previously published Expression of Concern. The original article stands as published, and we affirm the integrity of the research.

Enas Tariq Khudair, Onsa Lazzez, M. Zaied et al. · 0 citations
Open access Jul 2026

Automated Summarization Tool

The design realization and evaluation of an Automated Summarization Tool (AST) is presented which is a document intelligence platform based on google gemini 2.5 flash that outperforms the strongest fine-tuned transformer baselines (PEGASUS, BART) by ~14 points and is clearly ahead of BERTSUM-ext (a strong transformer baseline), Pointer-Generator Network, TextRank.

K. Kumar, A. Amandeep, Dharmender Kumar et al. · 0 citations