Skip to content
Review Open access

A Survey of Zero-Shot Sensitive Information Detection Techniques based on Large Language Models

Unknown authors
Aug 2026 · Scientific Journal of Intelligent Systems Research · 0 citations · 24 references

Abstract

With the rapid growth of digital information, the risk of sensitive information leakage in textual data, including personally identifiable information, medical privacy, financial data, and corporate confidential information, has become increasingly prominent. Traditional sensitive information detection methods, which mainly rely on rule matching, supervised learning, and manual annotation, struggle to meet the requirements of identifying diverse, open-domain, and dynamically evolving sensitive information. In recent years, Large Language Models (LLMs) have provided a new technical paradigm for zero-shot sensitive information detection without annotated data, owing to their powerful semantic understanding, contextual reasoning, and knowledge transfer capabilities. Through prompt learning, in-context learning, and instruction-driven information extraction approaches, LLMs can achieve flexible sensitive information identification in scenarios involving unknown sensitive categories and cross-domain applications. However, LLMs themselves introduce new security risks, including training data leakage, privacy memorization, and prompt injection attacks, posing significant challenges to sensitive information detection technologies. This paper presents a systematic survey of the development of LLM-based zero-shot sensitive information detection techniques. First, it introduces the development path of sensitive data detection technology, including rule-based methods, older machine-learning techniques, and currently popular pre-trained language models. Then it introduces the research content of LLM-driven zero-shot named entity recognition, open information extraction and privacy detection. List the privacy-leakage risks and corresponding defence measures for current applications of LLMs. Finally, this paper presents some future research directions for the above work and provides a path for the development of efficient, secure and trustworthy intelligent sensitive information detection systems.

Read PDF