Knowledge Accuracy and Response Characteristics of an On-Device Large Language Model in Building Environmental Engineering: a Preliminary Case Study of EXAONE 4.0 1.2B Using ChatGPT-4o as a Cloud Reference
Aug 2026· Architectural research· Vol 28· 0 citations· 30 references
Abstract
This preliminary case study examines EXAONE 4.0 1.2B, an on-device large language model (LLM), in building environmental engineering, using ChatGPT-4o as a high-capability cloud reference rather than a size-matched competitor. Ten Korean-language items were evaluated: four multiple-choice, three single-step calculations, and three short-answer questions covering indoor environmental quality, energy, and post-occupancy evaluation. The same simple prompt condition was applied to both models, with no model-specific prompt optimization; therefore, the results represent one standardized condition rather than each model’s maximum performance. Both models answered all seven objective items correctly, although EXAONE showed terminology confusion in its reasoning on a thermal-environment item. Five doctoral-level experts rated the short-answer responses for accuracy, completeness, and logical consistency. Mean scores were 4.18 for ChatGPT-4o and 2.98 for EXAONE. Fleiss’ kappa was 0.020 and 0.044, while free-marginal kappa was 0.333 and 0.222, respectively; the latter values indicate fair, not moderate, agreement, so absolute expert scores require cautious interpretation. Qualitatively, ChatGPT-4o more consistently organized responses around concepts and scope, whereas EXAONE tended to provide specific technical, operational, and Korean-regulatory content but occasionally omitted canonical elements or added unsupported detail. Because the item pool is small and only one on-device model was tested, the findings are exploratory and cannot be generalized to building environmental engineering as a whole or to on-device LLMs as a class. The results support further investigation of on-device LLMs as supervised offline assistants where connectivity or external data transmission is constrained, but not as unsupervised tools for engineering calculation or compliance decisions.
The proliferation of large language models (LLMs) across enterprise, research, and public-sector applications has created an urgent need for rigorous, multi-dimensional evaluation frameworks. This paper presents a comprehensive comparative analysis of seven state-of-the-art LLMs — GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro and Flash, LLaMA 3 70B, Mistral Large, and Claude 3 Haiku — across eight evaluation dimensions: benchmark accuracy, safety alignment, cost efficiency, inference latency, context handling, deployment flexibility, multilingual capability, and scalability. A weighted Multi-Criteria Decision Analysis (MCDA) framework is applied to produce transparent composite rankings from empirical benchmark data using five standardized benchmarks (MMLU, HumanEval, HellaSwag, GSM8K, MATH). Results indicate that Claude 3.5 Sonnet achieves the highest MCDA composite score (0.801), driven by accuracy (90.4% MMLU, 92.0% HumanEval) and safety alignment (4.9/5). Gemini 1.5 Flash emerges as optimal for cost-sensitive deployments ($0.075/1M tokens; 210 tok/s). The paper analyzes architectural trade-offs between dense transformers and Mixture-of-Experts designs, provides a deployment recommendation matrix, and contributes an extensible, evidence-based decision framework for enterprise AI practitioners.
Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.
Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili et al.· 0 citations
This work presents the first cross-task empirical evaluation of LLMs spanning five RE-related activities, as well as replication materials supporting reproducibility, and a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.
Jacek Dabrowski, Manjeshwar Aniruddh Mallya, Alessio Ferrari et al.· 0 citations
Penetration testing reports are a critical artifact in the cybersecurity workflow, yet their technical complexity frequently limits their utility for non-specialist stakeholders involved in risk remediation decisions. This paper investigates the feasibility of using four open-weight large language models, DeepSeek-r1:32b, Qwen3.5:35b, Gemma4:31b, and GLM-4.7-flash:32b, to generate plain-language summaries of penetration testing reports. A corpus of 65 publicly available pentest reports was used for evaluation. Model outputs were assessed across four dimensions: readability, technical term density, semantic similarity to the source document, and factual correctness using LLM-as-a-judge evaluation. Two classical extractive methods, LSA and TextRank, were included as baselines. Readability analysis using seven established metrics showed that Qwen3.5 and Gemma4 produced the most accessible summaries, reducing mean Flesch Reading Ease scores from 25.2 in the originals to 49.2 and 52.1 respectively, and lowering grade-level scores from post-graduate to high-school equivalents. Results across the remaining evaluation dimensions further indicate that appropriately selected open-weight LLMs can produce accessible and factually grounded summaries of technical security documents, offering a practical alternative to proprietary solutions in privacy-sensitive deployment contexts.
Prerit Datta, M. Islam, Ryan Wojciechowski· Annual International Compute...· 0 citations
The Construction Industry Institute (CII) has conducted extensive research over the past four decades, culminating in the identification of 17 best practices (BPs) proven to enhance overall project success when effectively implemented. However, the BPs remains underutilized in the construction industry due to the vastness of the related knowledgebase, heterogeneous data formats, and time-intensive nature of interpretation. This research addresses these challenges by developing a new CII BP handbook and an artificial intelligence (AI) tool designed to improve CII knowledge accessibility and usability for member companies. The first tool,
BP Primer
, was developed using a hybrid approach that combines qualitative analysis with retrieval-augmented generation (RAG)–based question and +answering. It extracts actionable insights (termed as “golden nuggets”) from 54 CII research reports and organizes them across the 17 BPs and project lifecycle phases, enabling users to quickly locate strategies tailored to their specific needs. Additional features include BP-level and report-level executive summaries and frequently asked questions (FAQs) to support high-level CII BP comprehension. The second tool, a
BP Conversational AI
, is built on a multimodal RAG framework, allowing users to interact with the BP multimodal knowledge base through natural language queries and document-grounded responses. Both tools were piloted with industry partners, and key lessons learned from the implementation have been documented. This research establishes a foundation for applying multimodal large language model (MLLM)-driven frameworks to construction knowledge curation. The curated knowledge can effectively guide construction professionals toward fit-for-purpose insights that support informed decision-making.
Gong Chen, M. Pedraza, Roya Albaloul et al.· Journal of Management in Eng...· 0 citations
This study aims to explore the potential of general-purpose Large Language Models (LLMs) in generating User Experience (UX) evaluations during the product verification phase, addressing issues in traditional UX research methods such as difficulties in user recruitment and long scheduling cycles. Using the User Experience Honeycomb model as the theoretical framework, the research selects the off-the-shelf GPT-4o as the experimental model. By combining optimized prompt engineering with multimodal inputs, a comparative analysis is conducted on the similarities and differences between LLMs and human users regarding evaluation coverage rate, language style, and problem perspectives. The experiment employs a deductive-inductive approach to code and analyze the collected evaluation data. The results indicate that the thematic overlap rate between LLM-generated evaluations and human user evaluations reaches 81.05%, demonstrating significant potential in simulating human users to output experience evaluations. Textual analysis reveals that LLM-generated UX evaluations exhibit strengths in systematic analysis, professional expression, and proactive risk identification; however, they show limitations in capturing nuanced emotions and dynamic interaction details. Additionally, the efficiency of UX evaluation is improved by 88.0% compared to human users. The study recommends adopting a Hybrid Intelligence evaluation model, leveraging the systematic analysis capabilities of LLMs while incorporating human users' acute perception of emotions and immediate experiences to enhance both the efficiency and comprehensiveness of UX research.
Xiaoyue Mao, Jun Zhang, Yijing Yang et al.· AHFE International· 0 citations