This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.
Abstract
Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such indicators is challenging, as relevant values may be dispersed or transformed within code. Although large language models (LLMs) have shown promise in security analysis, their ability to recover IOCs from malicious scripts remains underexplored. We present SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts. The benchmark comprises 634 manually verified JavaScript, PowerShell, and VBScript malware samples covering four IOC types (URLs, domains, IP addresses, and filesystem artifacts). We further stratify ground-truth IOCs by recovery level, distinguishing directly exposed indicators from those requiring decoding or reconstruction. Using this benchmark, we evaluate a broad range of proprietary and open-weight LLMs and show that IOC recovery without execution remains challenging across model scales: the strongest model reaches only 65.4 F1. To characterize how recovery fails, we introduce a false-positive taxonomy and use it to compare the error profiles of the evaluated models. We further study two mitigations on a small open-weight model, deterministic string utilities and task-specific adaptation, finding that they provide complementary recovery gains, raise precision, and shift errors toward sample-grounded mismatches.
The application of LLMs for detecting malicious PowerShell scripts and producing human-interpretable explanations for their classification decisions are investigated, showing that LLMs are capable of identifying and explaining malicious PowerShell scripts, although performance varies across different models.
Meng Wang, Emma Topolovec, B. Arana et al.· 0 citations
Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-b...
H. Choi, Wonyoung Jung, Haehoon Seo et al.· 0 citations
Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without c...
Arastoo Zibaeirad, Rodrigo Pato Nogueira, Marco Vieira· IEEE Transactions on Reliabi...· 0 citations
As sophisticated evasion techniques like polymorphism and staged execution increasingly neutralize conventional signature-based defenses, dynamic API sequence analysis has emerged as an effective approach for malware detection. However, extracting actionable intelligence from noisy execution logs while maintaining mode...
Dat Quoc Phan, Tien Duc Anh Hao, Nghi Hoang Khoa et al.· International Conference on...· 0 citations
The rapid evolution of malware variants has increasingly undermined traditional signature‐based detection techniques, which are easily evaded through obfuscation and polymorphism that preserve malicious functionality. This challenge is particularly acute in Internet of Things (IoT) environments, where device heteroge...
Khizar Hayat, Sadaf Hina, Fabiha Hashmat et al.· Security and Privacy· 0 citations
Python package security is largely source-centric, yet Python runtimes can execute bytecode directly through .pyc files, compiled-only modules, and marshalled code objects, creating an inspection-execution gap. We present an empirical study of Python bytecode as a security artifact. We measure bytecode exposure in PyPI...