Detecting Context-Dependent Sensitive Data in Unstructured Text
Abstract
The massive amount of publicly available data has necessitated an increase in public and organizational awareness of the potential risks of leaking private data, whether intentionally or unintentionally. The damage caused by leaking these data depends on their degree of sensitivity. Disclosing a person’s or an organization’s private data via different social media platforms might threaten people’s lives or the organization’s reputation or finances. Handling big data, especially unstructured data, is challenging. Consequentially, many solutions have been proposed to detect sensitive data in structured containers. However, detecting sensitive data in unstructured containers is still challenging, especially with context-dependent and high-performance measurement results. In this study, experiments on certain machine learning models and two transformers—DistilRoberta and ALBERT—were conducted to detect unstructured, textual, context-dependent sensitive data. The results show that DistilRoberta demonstrated higher accuracy and recall, and was faster and lighter than ALBERT.