Skip to content
Open access

REMEDy: a dataset for rationale extraction and span-based moderation of dialogue prompts

Aug 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 31 references

Abstract

The wide adoption of conversational AI systems necessitates urgent and interpretable safety moderation, especially given that Large Language Models (LLMs) continue to exhibit vulnerabilities despite alignment efforts, posing significant risks to individual users, organisations, and society. The ideal AI safety moderation system must be transparent and structurally interpretable. However, current moderation approaches typically rely on coarse classifications that offer limited interpretability and fail to capture the nuanced intent and contextual dependencies present in real-world user inputs. To advance moderation beyond these coarse labels, we present REMEDy, a novel dataset specifically built for extracting fine-grained rationales from user prompts. REMEDy features span-level annotations covering a broad taxonomy of safety-relevant categories, allowing for overlapping and nested textual spans to reflect complex prompt structures. Using REMEDy, we fine-tune multiple LLMs and evaluate their performance across two tasks: (i) rationale extraction, assessing their ability to accurately localise and classify harmful or ambiguous content; and (ii) prompt moderation, measuring improvements over state-of-the-art safety detectors. Our experiments demonstrate that REMEDy-trained models achieve competitive or superior moderation outcomes while simultaneously providing structured, human-readable rationales. REMEDy thus offers a valuable resource for developing safer, more transparent, and context-sensitive moderation systems.

Read PDF