Open access
Aug 2026
A Lightweight BERT-Variant Optimized for WebGPU-Based Real-TimeInference in Web Browsers
A lightweight BERT-inspired architecture optimized for GPU parallelmatrix operations through TensorFlow.js with the WebGPU backend is proposed, establishing a practical framework for deploying real-time edge NLP applications using open web standards and GPU acceleration.
Md Istiak Morsalin, Tasnim Akter Onisha, A. Shalan et al.
· Informatica · 0 citations