Evaluating the Efficiency of Open-Source AI Chat Models in Local Deployments
Abstract
This study evaluates the performance and operational practicality of locally deployed large language models (LLMs) as organizations seek cost-efficient, privacy-preserving alternatives to cloud AI. The objective is to generate workflow-aware evidence that supports informed model selection for automation use cases. Four freely available models, Minicpm-o-2_6, Gemma-3-12B-IT, Pixtral-12B, and Qwen2-VL-7B-Instruct, served as the experimental materials. Each model was deployed through LM Studio and integrated into n8n to simulate real-world automation pipelines, including notification flows. The method involved structured and repeatable testing across three input categories: text-only, text-and-image, and image-only. Automated logging captured response time, success rate, throughput, and system resource consumption (CPU, GPU, and memory). We also measured workflow-level overhead to assess the impact of orchestration on latency. The results show clear trade-offs between inference speed, resource consumption, and workflow-level efficiency. Minicpm-o-2_6 demonstrated the best overall balance with low latency (~5 seconds) and minimal CPU and memory usage. Qwen2-VL-7B-Instruct achieved the fastest average response time (~4 seconds) but required higher GPU and memory resources, positioning it for speed-critical workloads. Pixtral-12B delivered stable results with efficient GPU utilization yet slower inference. Gemma-3-12B-IT exhibited the highest latency (~17 seconds) with only moderate resource demands, making it the least efficient in workflow contexts. Across all models, end-to-end workflow latency (~33 seconds) exceeded inference time due to orchestration overhead. Multimodal inputs further increased latency by up to 300%. These findings highlight that workflow orchestration, rather than model inference alone, is the primary determinant of real-world latency in local automation pipelines.