Skip to content

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

Sep 2026 · 0 citations · 67 references
Computer Science

TL;DR

This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.

Abstract

Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such indicators is challenging, as relevant values may be dispersed or transformed within code. Although large language models (LLMs) have shown promise in security analysis, their ability to recover IOCs from malicious scripts remains underexplored. We present SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts. The benchmark comprises 634 manually verified JavaScript, PowerShell, and VBScript malware samples covering four IOC types (URLs, domains, IP addresses, and filesystem artifacts). We further stratify ground-truth IOCs by recovery level, distinguishing directly exposed indicators from those requiring decoding or reconstruction. Using this benchmark, we evaluate a broad range of proprietary and open-weight LLMs and show that IOC recovery without execution remains challenging across model scales: the strongest model reaches only 65.4 F1. To characterize how recovery fails, we introduce a false-positive taxonomy and use it to compare the error profiles of the evaluated models. We further study two mitigations on a small open-weight model, deterministic string utilities and task-specific adaptation, finding that they provide complementary recovery gains, raise precision, and shift errors toward sample-grounded mismatches.

View source

Similar papers

Detection and Explanation of PowerShell Malware with Large Language Models

The application of LLMs for detecting malicious PowerShell scripts and producing human-interpretable explanations for their classification decisions are investigated, showing that LLMs are capable of identifying and explaining malicious PowerShell scripts, although performance varies across different models.

Meng Wang, Emma Topolovec, B. Arana et al. · 0 citations
#machine learning Preprint Sep 2026

ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers

Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-b...

H. Choi, Wonyoung Jung, Haehoon Seo et al. · 0 citations
2026

LLMKernelBench: Benchmarking Large Language Models on Software Vulnerability Detection in Linux Kernel

Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without c...

Arastoo Zibaeirad, Rodrigo Pato Nogueira, Marco Vieira · 0 citations
Conference Aug 2026

Explainable Malware Detection from Noisy API Sequences with RAG-Based MITRE ATT&CK Mapping

As sophisticated evasion techniques like polymorphism and staged execution increasingly neutralize conventional signature-based defenses, dynamic API sequence analysis has emerged as an effective approach for malware detection. However, extracting actionable intelligence from noisy execution logs while maintaining mode...

Dat Quoc Phan, Tien Duc Anh Hao, Nghi Hoang Khoa et al. · 0 citations
Open access Sep 2026

Base Semantics—A Novel Approach for Minimizing Malware Detection Rules

The rapid evolution of malware variants has increasingly undermined traditional signature‐based detection techniques, which are easily evaded through obfuscation and polymorphism that preserve malicious functionality. This challenge is particularly acute in Internet of Things (IoT) environments, where device heteroge...

Khizar Hayat, Sadaf Hina, Fabiha Hashmat et al. · 0 citations
Preprint Aug 2026

Beyond Source: An Empirical Study of Python Bytecode Security Risks

Python package security is largely source-centric, yet Python runtimes can execute bytecode directly through .pyc files, compiled-only modules, and marshalled code objects, creating an inspection-execution gap. We present an empirical study of Python bytecode as a security artifact. We measure bytecode exposure in PyPI...

Bai-Hong Chen, Tian Xie, Wen Li · 1 citation · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.