Explainable Prompt Injection Detection using Sentence Embeddings, Random Forest, and Word-Level Attribution
A explainable framework that detects the attacks using a Random Forest (RF) classifier along with semantic sentence embedding and SHAP (Shapley Additive Explanations) is used together with a wordlevel attribution mechanism to give human-legible explanations for the model predictions, and to emphasize harmful parts within the prompt.