Skip to content

Category

small language model

709 papers

#machine learning Preprint Aug 2026

Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization

Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. In this work, we identify that the evolving embedding structures are closely related to token-conditioned label and contextual distributions, which we formalize as probability signatures. We observe a progressive learning process, which we term Context Staircase: embeddings learn the low-order statistic signatures of the data before the high-order ones. More specifically, we observe that early in training they align with the simplest, context-free signature linking a token to its label, and as training proceeds, they progressively reflect signatures involving more and more context tokens. We then analyze the gradient flow of embeddings under small initialization to explain this phenomenon, deriving embedding evolution equations for feed-forward and self-attention architectures. We further extend these observations to real language-model training. Finally, we show that these embedding structures play an important role in both task learning and the incorporation of semantic structure into the embedding space. Overall, our results provide a dynamic explanation of how data statistics and architecture jointly shape token embeddings in language models, and reveal an implicit bias in the space of data statistics: training proceeds from simpler, low-order statistical relations toward increasingly complex, context-dependent ones.

Junjie Yao, Liangkai Hang, Zhi-Qin John Xu · 0 citations
#machine learning Preprint Aug 2026

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confidence retention (L), and grounding/relation-probe instability (A). Across 58K+ examples from four benchmarks and four MLLMs, image-removal confidence retention is most prevalent, while grounding/relation-probe instability better separates failure families. Only 18 of 48 source-target checks are diagonally aligned, so the coordinates should be interpreted jointly rather than as independent causal sources. With the dataset fixed, the joint signature improves failure-family AUROC from 0.634 to 0.769 on HallusionBench and from 0.707 to 0.817 on VizWiz, with smaller gains on POPE and VSR. In pooled XGBoost analysis, AUROC rises from 0.78 with scalar confidence to 0.95 with (V, L, A) and 0.97 when confidence is added. The same signature does not automatically improve correctness ranking. The three tested direct scalarizations can harm it. These results separate failure diagnosis from abstention scoring: multimodal uncertainty should characterize failure structure before it is used to decide whether to abstain or correct.

Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty · 0 citations

An interpretable river water quality prediction model by integrating deep learning and large language models

Physical mechanism-based models for river water quality prediction involve complicated calculations,whereas machine learning models lack interpretability,resulting in a disconnection between predictions and management decisions that hinders practical application. To deeply integrate high-precision prediction with decision support,a TCN-Attention deep learning model that combines a temporal convolutional network and an attention mechanism is constructed to predict six core water quality indicators,with Bayesian optimization used for automatic parameter tuning. The SHAP interpretability technique is introduced to quantify the contributions of multi-source input features and reveal key driving factors. Based on the Qwen3-Next large language model,water quality grade evaluations and improvement recommendations are automatically generated. Application results for rivers in Xinwu District of Wuxi City demonstrate that the TCN-Attention model achieves good prediction performance under small-sample conditions,with domestic water use,turbidity,and water identified as the most critical influencing factors. Qwen3-Next model achieves an accuracy of 89.8% in water quality grading,and the proposed improvement recommendations are targeted and practicable. The proposed interpretable intelligent water quality prediction method effectively improves the accuracy and interpretability of urban river water quality predictions,providing a reliable technical pathway for smart water environment management.

Zuxiang Situ, Hongwu TANG, Qihua Ran et al. · 0 citations
#computer vision Conference Jan 2003

Experimental software engineering (STESE)

Software engineering theory and practice is still to a large extent based more on faith than on science. Only by contributing to the scientific and empirically grounded body of knowledge within a specific area of application, theory and practice can develop. Experimentation is an important scientific approach to collect empirical data and to test theories as well as to bring light to new phenomena so that theories can be formulated and corrected. This is the background for the emerging field of experimental software engineering. The focus of this minitrack is on experiments and experimental studies performed in academic or industrial settings where the aim is to study the software professionals' work practices related to the development of software. This minitrack is divided in two three-paper sessions. The papers are briefly introduced in the following. The three papers in the first session are experiments performed in an academic setting. Syversen, Anda and Sjoberg report the results from an experiment with 26 subjects where they explore how a use case model can best be applied in an object-oriented development process. Serrano, Calero and Piattini describe how to apply the experimental method in metrics definition for multidimensional data models. Their paper gives an overview of the method including a description of how it was applied. The first session is concluded with a paper authored by Liu and Grandon where they empirically explore with 79 subjects how task performance and domain-specific self-efficacy influence the perceived ease of use of object-oriented analysis techniques. The first two papers in the second session include a set of experiments and an empirical study performed in an industrial setting. Jokela describes five different experiments where the attempt is to assess the quality of the usability engineering processes of four different companies. Jokela explains how the assessment process is iteratively changed and improved based on the results of the earlier experiments. Borjesson and Mathiassen compare two software process improvement initiatives carried out in industry. They focus on factors affecting the implementation success. Dugan, Glinert and Rogers conclude the minitrack by introducing a technology-focused methodology called CAMELOT, which is intended for testing computer supported co-operative work software. They report results from an experiment where the proposed methodology was tried out.

K. Kautz, P. Abrahamsson · 1 citation
#computer vision Open access Sep 1999

Commitment to Software Process Improvement—Development of Diagnostic Tool to Facilitate Improvement1

This paper suggests that by operationalizing the concept of commitment in the shape of a model, a new insight is provided in improving software processes—a more human centered approach as opposed to various technical approaches available. In doing so the SPI managers/change agents are able to plan better the software process improvement initiative and benchmark successful projects (as well as failed ones). Results from five interviews with SPI professionals on the proposed Behavior-based Commitment Model are reported, together with early results from the empirical test in 14 software process improvement projects. Early results suggest that the behaviors introduced in the model are relevant in SPI initiatives, the use of model raises the awareness about the people issues in improving processes, and the model could be used aside with CMM, SPICE or other process improvement models.

P. Abrahamsson · 10 citations
#computer vision Open access Jun 2000

Modelling Usability Capability - Introducing the Dimensions

Usability capability is a characteristic of a development organization that predicts the level of usability the development projects are capable of achieving. Our experiments with the existing usability capability models indicate that current process assessment methods do not discover all relevant problems that might impede effective user-centered design (UCD) in development organizations. We propose an enhanced model where the usability capability is analyzed from three dimensions: user-centered infrastructure, implementation of user-centered practices in development projects, and business management commitment to usability as a competitive asset.

Timo Jokela, P. Abrahamsson · 14 citations
#computer vision Conference Open access Dec 2002

The personal software process: experiences from Denmark

The focus of the research and practice in software process improvement (SPI) is shifting from traditional large-scale assessment based improvement initiatives to smaller sized, tailored initiatives where the emphasis is on the development personnel and their personal abilities. Personal software process (PSP/sup SM/) is a method designed for improving the personal capabilities of the individual software engineer. This paper contributes to the body of knowledge within this area by reporting experiences from Denmark. The findings indicate an improvement in effort estimation skills and an increase in the resulting product quality in terms of reduced total defect density. The data shows that even with a relatively small effort (i.e., 10%) used in defect prevention activities (i.e., design and code reviews) almost one third of all defects could be removed and, consequently, the time required for the testing was reduced by 50%. On the basis of this data, the use of the PSP method in the software industry is discussed.

P. Abrahamsson, K. Kautz · 17 citations · ⚡2
#computer vision Open access Dec 2002

Commitment Nets in Software Process Improvement

This study suggests that software organizations operate through strategic, operational and personal commitment nets, and shows that SPI is driven through the formation and reformation of commitment nets.

P. Abrahamsson · 35 citations · ⚡2
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new angle and explores waste in the Kanban-driven software development project context. A preliminary research model is presented for helping the consequent replication of the study. The results from the empirical analysis suggest Kanban can be an effective method in visualizing and organizing the current work, but does not prevent waste from creeping in, although the overall project outcome may be successful.

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.