Skip to content

Category

small language model

845 papers

The role of artificial intelligence in developing meal plans in outpatient dietetics: A feasibility and proof of concept study.

BACKGROUND Personalized meal planning by registered dietitian nutritionists (RDNs) is time-intensive. Large language models (LLMs) may automate drafting meal plans, but their nutritional accuracy in clinical practice is uncertain. METHODS In this proof-of-concept study, five outpatient RDNs and four LLMs (Gemini, CoPilot, ChatGPT 4.0, and customized ChatGPT 4.0) each generated 3-day meal plans for five validated clinical scenarios. Effectiveness was defined as accuracy in meeting pre-specified energy, protein, carbohydrate, fat, and sodium targets. Time to create plans and RDN comfort (self-rated confidence in nutritional accuracy and clinical appropriateness on 1-5 Likert scale) were recorded. Three independent RDNs, blinded to source, analyzed nutrient content using Nutritionist Pro. Group differences were assessed with t-test and ANOVA. RESULTS All LLMs and RDNs produced feasible meal plans. LLMs generated meal plans in under 1 min, whereas RDNs required a mean of 44 min per scenario. RDNs reported comfort levels ranging from 3.8 to 4.8. Across most scenarios, LLM plans delivered a smaller proportion of requested energy than RDN plans, which more consistently approached energy targets. Both groups performed similarly for the Mediterranean diet scenario. Overall, protein accuracy did not differ. However, in chronic kidney disease, LLMs undershot the guideline-based protein target, while RDNs tended to modestly exceed it. Accuracy for low-carbohydrate, fat, and sodium diets was comparable. CONCLUSION LLMs can rapidly generate clinically plausible meal plans but are less reliable than RDNs in achieving prescribed energy and selected macronutrient goals. Prompt precision is essential for nutrient-specific targets. A hybrid model in which RDNs refine LLM-generated drafts may leverage efficiency without sacrificing clinical accuracy.

M. Mundi, Osman Mohamed Elfadil, Danielle P. Johnson et al. · 0 citations
#small language model Preprint Aug 2026

FlashAttention for Scalable Vector Architectures

This paper presents FlashAttention-V, a blocked FlashAttention for scalable vector architectures that adapts efficiently from short to very long vectors by exploiting parallelism across attention heads, inter-head packing to enable efficient utilization of vector lengths beyond the head dimension, and improving vector register utilization and memory access locality.

Sonia Rani Gupta, Nikela Papadopoulou, Miquel Pericàs · 0 citations
#small language model Preprint Aug 2026

COSTA: A Cluster-Centric Paradigm for Annotation-Free Open-Set Semantic Segmentation of Aerial Point Clouds with Domain Shifts

COSTA leverages the domain gap through proven test-time adaptation, and groups each batch of target-domain points into a small set of semantic clusters based on the similarity distribution in the adapted feature space, and propagates high-confidence pseudo labels obtained from an open-vocabulary vision-language model to all points through cluster-level voting.

Yanghong Lin, Li Fang, Tianyu Li et al. · 0 citations
#small language model Review Aug 2026

Large Language Models in Oral and Maxillofacial Surgery Triage: A Scoping Review

Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.

K. Surendran, I. Aziz, Glyndwr Jenkins · 0 citations
#small language model Preprint Aug 2026

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

Matteo Cargnelutti, Catherine Brobston, Eben English et al. · 0 citations
#small language model Preprint Aug 2026

Mise-en-Sc\`ene: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

The designs produced by Mise-en-Sc\`ene are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while the match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.

Zipeng Xu, Ryan Murdock, Umberto Michieli · 0 citations
#small language model Review Open access Aug 2026

Business model innovation in SMEs for sustainable and digital transformation: a systematic review and integrative framework

Small- and medium-sized enterprises (SMEs) increasingly need to reconfigure how they create, deliver, and capture value in response to digitalisation, sustainability demands, and environmental uncertainty. This study systematically reviews 162 English-language articles published in SSCI-indexed journals to integrate fragmented research on business model innovation (BMI) in SMEs. Descriptive analysis, keyword co-occurrence analysis, and article-level qualitative content analysis are combined. The findings identify four broad knowledge domains and show that 41 studies treat BMI primarily as an outcome, 44 as an organisational process, 17 as an antecedent of subsequent outcomes, and 36 as part of a combined causal relationship; a further 24 studies are descriptive or typological. Across these roles, the literature explains BMI through resource mobilisation, entrepreneurial action logics, absorptive and dynamic capabilities, experimentation, and network mobilisation. Its effects on performance, growth, resilience, internationalisation, and sustainability depend on complementary resources, implementation capabilities, stakeholder alignment, and environmental conditions. The evidence supports an adaptive and iterative interpretation of SME BMI, but direct evidence of recursive feedback remains limited. Resource constraints also have conditional effects, stimulating bricolage and focused experimentation in some circumstances while inhibiting substantive change when shortages become severe. The review provides an integrative synthesis specific to SMEs and advances digital business model innovation research by explaining how digital pressures are translated into changes in value proposition, value creation and delivery, and value capture through organisational and relational mechanisms.

Bi Zhang, Nurul Atasha Jamaludin, Shanshan Yue et al. · 0 citations
#small language model Review Aug 2026

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making.

M. Zamani, Fatemeh Ziaeetabar · 0 citations
#small language model Preprint Aug 2026

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurements take the form y = Ax + w, where A describes the interventions, x is the target map, and w contains nonlinear response, sampling error, and basis misspecification. This language gives a practical procedure: start with the least costly measurements, test on held-out interventions at the intended scale, calibrate simple mismatch, and expand the measurement family when structured residuals remain. Control provides a demanding validation setting because an estimate that guides an intervention acts as an observer. In a two-HMM model, control error rises with observer error, while target improvement can hide nuisance-state movement. Under forward-only access, sparse aggregate measurements recover a finite-effect map with fewer interventions than coordinate patching. With gradient access, finite probes improve a local attribution map. Lifted measurements and Hessian-vector products recover interactions missed by first-order maps, while Tracr shows that the required family depends on the basis. On GPT-2-small IOI, the Name Mover-Negative Name Mover interaction is the largest held-out predictive term among three tested cross-group pairs. On Qwen-2.5-7B, finite calibration makes an additive refusal-response map adequate, so held-out error does not support pairwise lifting.

Vijay Erramilli · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.