Skip to content

EVALUASI KEAMANAN PROMPT DALAM INTERAKSI MANUSIA DAN AI: STUDI KASUS PADA MODEL GPT-4

Sep 2026 · ACADEMIA Jurnal Inovasi Riset Akademik

Abstract

ABSTRACT The rapid development of generative artificial intelligence, particularly language models such as GPT-4, has transformed human–machine interaction through the use of prompts across various contexts. However, the adaptive capabilities of these models introduce security challenges, particularly those involving prompt injection, output manipulation, behavioral constraint bypassing, and potential data breaches. This study aims to analyze the security dimensions of prompt-based interaction with GPT-4 and to identify its vulnerabilities and the effectiveness of the implemented mitigation mechanisms. A case study method was employed through several stages, including the design and testing of various types of manipulative prompts, observation of model responses, identification of deviations from security constraints, and evaluation of the available security filters and behavioral restrictions. The findings indicate that security filters and behavioral constraints implemented in GPT-4 can reduce certain risks associated with harmful interactions, but they do not completely prevent more sophisticated prompt-engineering strategies. These vulnerabilities demonstrate that the security of language-model-based systems cannot rely solely on built-in protection mechanisms but requires layered risk mitigation strategies, continuous security evaluation, and improved user awareness. The study concludes that strengthening security frameworks, conducting systematic testing against manipulative scenarios, and promoting safe and ethical AI usage practices are necessary to improve the resilience of prompt-based AI systems. ABSTRAK Perkembangan kecerdasan buatan generatif, khususnya model bahasa seperti GPT-4, telah mengubah pola interaksi manusia dan mesin melalui penggunaan perintah (prompt) dalam berbagai konteks. Namun, kemampuan model dalam memahami dan menghasilkan respons secara adaptif menimbulkan tantangan keamanan, terutama terkait injeksi perintah, manipulasi keluaran, pengabaian batasan perilaku, dan potensi pelanggaran data. Penelitian ini bertujuan menganalisis dimensi keamanan interaksi berbasis perintah pada GPT-4 serta mengidentifikasi kerentanan dan efektivitas mekanisme mitigasi yang diterapkan. Penelitian menggunakan metode studi kasus dengan tahapan berupa perancangan dan pengujian berbagai jenis perintah manipulatif, pengamatan terhadap respons model, identifikasi bentuk penyimpangan dari batasan keamanan, serta evaluasi terhadap filter dan mekanisme pembatasan perilaku yang tersedia. Hasil penelitian menunjukkan bahwa penerapan filter keamanan dan pembatasan perilaku pada GPT-4 mampu mengurangi sebagian risiko interaksi berbahaya, tetapi belum sepenuhnya mencegah strategi rekayasa perintah yang lebih kompleks. Kerentanan tersebut menunjukkan bahwa keamanan sistem berbasis model bahasa tidak hanya bergantung pada mekanisme perlindungan bawaan, tetapi juga memerlukan strategi mitigasi berlapis, evaluasi keamanan secara berkelanjutan, dan peningkatan kesadaran pengguna. Penelitian ini menyimpulkan bahwa penguatan kerangka kerja keamanan, pengujian terhadap skenario manipulatif, dan penerapan praktik penggunaan AI yang aman dan etis diperlukan untuk meningkatkan ketahanan sistem AI berbasis perintah.

Read PDF

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4
#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3

Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNets

This paper designs Markov decision processes (MDPs) for different combinatorial problems and proposes to train conditional GFlowNets to sample from the solution space and demonstrates that GFlowNet policies can efficiently find high-quality solutions.

Dinghuai Zhang, H. Dai, Esmeralda S. Whitammer et al. · 59 citations · ⚡8

Ethically Aligned Design of Autonomous Systems: Industry viewpoint and an empirical study

An empirical study on the current state of practice in artificial intelligence ethics is conducted by means of a multiple case study of five case companies, which indicates a gap between research and practice in the area.

Ville Vakkuri, Kai-Kristian Kemell, Joni Kultanen et al. · 56 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.