Skip to content

Author

P. Buitelaar

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Leveraging Large Language Models to Detect and Revise Unsafe Responses in Context-Sensitive Dialogues

Large Language Models (LLMs) excel at tasks like classification, summarisation, question answering among others, with performance comparable to humans. Despite these capabilities, leveraging LLMs to transform unsafe responses in context-sensitive dialogues is underexplored. In this work, we propose a pipeline that leverage LLMs as safety detector, editor and evaluator to mitigate undesired behaviour in human-computer dialogues. At the first iteration, our experimental results on two evaluation datasets show reduction in the unsafe dialogues from 47% to 13% and 48% to 2% respectively, with 82% and 92% agreement between the safety detector and evaluator after revision. Human evaluation of randomly sampled dialogues demonstrates reduction in unsafe responses after revision. Additionally, the revision LLM (editor) exhibits a higher proportion of refusals without compromising fluency and coherence of the revised dialogues.

T. Ajayi, M. Arcan, P. Buitelaar · 0 citations