Skip to content
#software testing Open access

Code generation for legal metadata extraction: a decomposition-based in-context learning approach

Sep 2026 · Requirements Engineering · Vol 31 · 0 citations · 62 references

TL;DR

The results demonstrate that, once a domain metamodel and expert-authored examples are available, few-shot code generation can extract legal metadata relationships without training a task-specific supervised model and can be adapted to unseen legislation.

Abstract

Software systems must comply with legal regulations, which is a resource-intensive task, particularly for small organizations and startups lacking dedicated legal expertise. Extracting metadata from regulations to elicit legal requirements for software is a critical step to ensure compliance. However, it is a cumbersome task due to the length and complex nature of legal text. Although prior work has pursued automated methods for extracting structural and semantic metadata from legal text, they do not consider the interplay and interrelationships among attributes associated with these metadata types, and they rely on manual labeling or heuristic-driven machine learning, which does not always generalize to new documents. In this paper, we introduce a decomposition-based in-context learning method for automatically generating a canonical representation of legal text encoded as executable Python code. Our representation is instantiated from a manually designed Python class structure that serves as a domain-specific metamodel, capturing both structural and semantic legal metadata and their interrelationships. Our corpus contains 13 US state data breach notification laws (332 paragraphs), of which six unseen laws (182 paragraphs) form the held-out test set. On this test set, our proposed method using GPT−5.1 achieves 90.5% semantic test accuracy with a precision of 79.4% and a recall of 81.9%. We also assess the generalizability of the method to the Children’s Online Privacy Protection Act (COPPA), a US federal law. The results demonstrate that, once a domain metamodel and expert-authored examples are available, few-shot code generation can extract legal metadata relationships without training a task-specific supervised model and can be adapted to unseen legislation.

Read PDF

Similar papers

Conference Open access Aug 2026

AI-Driven Knowledge Externalisation: From Unstructured Documents to Structured Data Models

The findings suggest that AI-based structured extraction may redefine how organisations formalise expertise, shifting from document-centric storage toward schema-driven knowledge architectures.

Dilyan Georgiev, E. Gourova · 0 citations

AI-Guided Metadata Construction for Meaning-Driven Digital Knowledge Systems: A Framework for Automated Metadata Generation and Semantic Discovery

An AI-guided framework is developed that aligns AI-assisted metadata extraction with Dublin Core Terms and the FAIR principles for digital libraries, archives, and cultural-heritage repositories and is evaluated as a design-science artefact in which retrieval is not a side feature but a feedback loop.

Wirapong Chansanam, Umawadee Detthamrong, Chunqiu Li et al. · 0 citations
#small language model Book Open access Aug 2026

Extracting Logical Structure in Code Documents via Semantic Segmentation and Language Models

Two language-model-based strategies are proposed for semantic code document segmentation, including a line-by-line approach that classifies each line of code separately before grouping the results into functional units, and a range-based approach that aims to directly determine groups of code lines from the input.

A. Dahou, A. Scherp, Sebastian Kurten et al. · 0 citations
Review Open access Aug 2026

NLP-Driven Extraction of Key Features from Legal Texts: Court Opinions, Briefs, Statutes, and Case Law

The proposed system aims to facilitate decision-making processes in legal practice and even the accuracy of the proposed model, which guarantees that accurate and reliable information is extracted and reduces the time and costs that conventionally come with manual legal analysis.

S. A. Gade, Sivaram Ponnusamy · 0 citations
Conference Aug 2026

Ontology-Guided Pedagogically Meaningful Knowledge Component Extraction from Code

Accurate extraction of Knowledge Components (KCs) is critical for fine-grained learner modeling in programming education. Yet existing approaches remain limited: manual Q-matrices ignore solution variability; Abstract Syntax Tree (AST) based methods may not produce pedagogically meaningful KCs; and Large Language Model...

Mathangi Krishnathasan, K. Hewagamage, E. Hettiarachchi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.