Skip to content
Open access

AI-Based Assistant for Generating and Analyzing Legal Documents

Jul 2026 · Turkish Journal of Engineering · Vol 10, pp. 758-771 · 0 citations · 10 references

TL;DR

This paper presents an AI-assisted tool based on a fine-tuned Distil-BERT model for effective classification of legal documents, semantic comprehension, and user-centric legal information support and demonstrates that distillation-based knowledge transformers can achieve better generalization with less complexity.

Abstract

The legal industry has been heavily impacted by fast-paced digitization, but the complete volume of complex legal documents, technical jargon, and arcane procedures continues to limit access to legal knowledge for both professionals and the general public. This highlights the need for more powerful, AI-based solutions for the retrieval and comprehension of legal information. The state-of-the-art AI and natural language processing (NLP)-based legal assistance systems are designed to increase accessibility, but are limited in their practical impact, scalability, and explainability. In this paper, we present an AI-assisted tool for the generation and analysis of legal documents. The system is based on a fine-tuned Distil-BERT model for effective classification of legal documents, semantic comprehension, and user-centric legal information support. The framework fulfils data preprocessing, transformer-based feature extraction, and supervised classification in an end-to-end manner, and can be embedded in legal service applications. The experimental results on a legal document dataset show that the proposed Distil-BERT model achieved an accuracy of 82.01% and a mean F1 of 84.32%. These results demonstrate that distillation-based knowledge transformers can achieve better generalization with less complexity, which is a desirable property for scalable solutions in the legal AI context. Further, we present important ethical and deployment factors to consider (e.g., data privacy, legal liability of AI outputs, and algorithmic bias mitigation). In sum, the proposed system takes a step toward intuition-friendly, efficient, and ethically conscious AI-supported legal aid that furthers the digital transformation of legal services.

Read PDF

Similar papers

Review Open access Aug 2026

NLP-Driven Extraction of Key Features from Legal Texts: Court Opinions, Briefs, Statutes, and Case Law

This paper outlines a unique method of legal text processing using Natural Language Processing (NLP) technology to extract the information from the legal texts meaningfully and naturally. The proposed system is designed in a data pipeline architecture by integrating the NLP functionalities such as tokenization, part-of-speech tagging, named entity recognition (NER), and dependency parsing to systematize the processing of typologies of legal text hubs, including legal briefs, statutes, and case law. The methodology presented concerns the importance of pre-processing legal texts that address domain-specific challenges. The texts may contain ambiguities, while the legal language itself is a very intricate kind of language. The system uses advanced methods like syntactic parsing and semantic role labeling to parse and find relevant entities, relationships, and context, ensuring the automation of the large amount of raw legal data for review and analysis. Besides that, the first is leveraging machine learning models to optimize the data extraction process and to ensure high efficiency and scalability. This methodology guarantees that accurate and reliable information is extracted and reduces the time and costs that conventionally come with manual legal analysis. The focal point of the offered system is overcoming legal workflow issues and bringing model texts to widespread use. Therefore, the proposed system aims to facilitate decision-making processes in legal practice and even the accuracy of the proposed model.

S. A. Gade, Sivaram Ponnusamy · 0 citations
Open access Jul 2026

RAG-Based AI Compliance Monitoring and Report Generation System

CompVault, an Enhanced Retrieval-Augmented Generation (ERAG)-based Artificial Intelligence Compliance Monitoring and Report Generation System for intelligent regulatory compliance assessment, and results indicate that the ERAG-based framework can be used as an efficient, scalable, and explainable solution for regulatory compliance monitoring and automated report generation.

S. N., Sathyapriya P., Vishnu Priya R. M. et al. · 0 citations
Open access Jul 2026

Law Buddy: An AI-Based Legal Information Assistant

AI Legal Buddy is an AI-powered legal information assistant designed to help users understand Indian laws in a simple and accessible way. The system allows users to enter their legal queries through text or voice input. The frontend is developed using React, TypeScript, Tailwind CSS, and Vite, providing a responsive and user-friendly interface. The backend is implemented using Supabase Edge Functions, which process user queries. In this project, we are using RAG (Retrieval- Augmented Generation) to retrieve relevant legal data from stored database and then generate the response based on that data. This helps in giving more accurate and meaningful answers instead of general responses. The system analyzes the query, identifies related Indian Acts and Sections, checks the seriousness of the issue, and provides structured responses with useful steps. The system also supports multiple languages to improve accessibility for users from different regions. By combining modern web technologies with RAG-based approach, the project aims to make legal information more understandable and easy to access for common people. Keywords: artificial intelligence; Retrieval-Augmented Generation; Legal Information Sys- tems; Multi-Agent Systems.

Tadiparti.Venu, Pakki Bhargava Shanmukha Sai, Lotheti Pravalika et al. · 0 citations
Conference Open access 2026

Developing and putting into practice an intelligent legal assistant for public procurement: An enhanced hybrid RAG method

The increasing complexity of public procurement regulations poses significant challenges for public administrations in accessing, interpreting, and managing regulatory information efficiently. In Morocco, this challenge is amplified by the large volume of legal documents, frequent regulatory updates, and the presence of numerous scanned archives. To address these issues, this paper proposes a sovereign Artificial Intelligence framework based on a Retrieval-Augmented Generation (RAG) architecture for regulatory knowledge management within public administration. The proposed solution operates entirely in an On-Premise environment, ensuring data confidentiality and digital sovereignty. It integrates an Optical Character Recognition (OCR) pipeline for processing scanned documents, a hybrid retrieval mechanism combining semantic and lexical search, and a locally deployed Large Language Model (Llama-3-8B-Instruct) for context-aware answer generation. The framework was evaluated using a corpus of more than 1,200 regulatory documents and a Golden Dataset composed of 50 expert-validated question-answer pairs. Experimental results achieved a Recall@5 of 92.4%, an MRR@10 of 0.94, and a semantic similarity score of 92.4%, outperforming conventional retrieval approaches while maintaining low response latency. The results demonstrate the potential of sovereign generative AI to enhance regulatory knowledge management, improve information accessibility, and support decision-making processes within Moroccan public administration.

Nihal Sajjaa, Hayat Bihri, Nasreddine Haqiq et al. · 0 citations
Open access Jul 2026

AI-Based Document Analysis and Question Answering System

This study provides an AI- Based document analyzer with a question-answer system that makes use of Natural Language Processing approaches that is affordable, scalable, and suitable for business, education, and research.

Radhika Sharma, Devraj Gautam · 0 citations