This work uses the statistically validated links obtained from the microcanonical bipartite configuration model as the reference benchmark for assessing performances of four widely used null models, and suggests that null models should be compared not only according to the constraints they preserve but also according to the statistical consequences that these constraints induce on the distribution of the test statistic.
Abstract
Statistical validation of projected bipartite networks depends critically on the null model adopted to describe random co-occurrences. Although several null models have been proposed, their comparison has mainly focused on the validated backbones they produce rather than on the statistical assumptions underlying their construction. Here we compare four widely used null models - the microcanonical configuration model generated by the Curveball algorithm, the Bipartite Configuration Model (BiCM), the Bipartite Partial Configuration Model (BiPCM), and the Hypergeometric approximation - using three empirical bipartite systems from comparative genomics, international trade and food science. We use the statistically validated links obtained from the microcanonical bipartite configuration model as the reference benchmark for assessing performances of the other three models. Our central result is that the statistical consequences of relaxing null-model constraints cannot be understood solely from the constraints themselves but must be analyzed through the probability distribution induced for the co-occurrence statistic. In particular, the combined behaviour of the expectation and variance largely explains the observed differences among the statistically validated backbones. We further derive a leading-order sparse approximation for the BiCM expectation, showing that the first correction to the Hypergeometric prediction is controlled by the degree heterogeneity of the non-projected layer. Surprisingly, despite neglecting this heterogeneity, the Hypergeometric model accurately reproduces the co-occurrence variance of the microcanonical ensemble across all datasets. Our results suggest that null models should be compared not only according to the constraints they preserve but also according to the statistical consequences that these constraints induce on the distribution of the test statistic.
This work proposes a novel and practical framework for dependence testing in labeled graphs via mutual information over a structure-weighted joint label distribution and demonstrates that the proposed test is a statistically sound and an effective tool for uncovering nontrivial dependencies in graph data.
Nikolaos Papagiannis, Vasam Manjveekar Prabantu, A. Grama et al.· Proceedings of the 32nd ACM...· 0 citations
Exponential random graph models (ERGMs) are now a staple for social network analysis due to their ability to parameterize effects from endogenous network structures, such as two stars and triangles. However, interpreting coefficients for structural ERGM terms poses two problems. First, it is not possible to change the value of a structural term “holding all else constant” in models with nontrivial dyadic dependencies. Second, parameters for structural ERGM terms may be affected by scaling (noncollapsibility), which can alter the direction, size, and significance of structural coefficients. While the first issue is known in the literature, the second has not been previously reported. This study introduces a methodological framework based on marginal structural effects (MSEs) to address both problems. MSEs capture the discrete marginal effect of an endogenous network structure by calculating the difference in probability when comparing two potential ties that differ only by the value of a structural term. MSEs provide an intuitive interpretation for structural network effects and are robust to the effects of scaling. Extensions to comparisons between models, interactions with exogenous covariates, and analysis of samples of networks are discussed. An example is provided using the largest AddHealth school network to demonstrate how the framework can be applied.
Scott W. Duxbury· Sociological Methods & R...· 0 citations
A rigorous statistical treatment of two-model comparisons on the same evals can be achieved by paired t-tests, analyzing their standard errors and a clustering correction for correlated questions. Nevertheless, leaderboards, ablations studies, and hyperparameter sweeps, usually compare $K>2$ models simultaneously. In this paper, we present a single random-effects model for scores on shared evals. We leverage models as a fixed effects and questions (or question cluster) as random effects. We show that fitting it a classical ANOVA or a linear mixed model we can recover Miller's paired and clustered estimators for K=2, but we extend the results to any K and to unbalanced, clustered designs. We validate the described model in a simulation study and using a real-data application. We take six openly available language models scored on 1,497 shared MMLU-Pro questions across 14 subject clusters, and we show that pairwise ranking claims survive depending on properly accounting for both question-level pairing and multiple comparisons. By the end of the manuscript, we provide concrete recommendations for reporting multi-model eval results.
The definition of directed graph wide-sense stationarity is revisited, and the surrogate signals preserve covariance under the stationary assumption to demonstrate the feasibility of the scheme to detect irregular node covariance and benchmark the method against conventional schemes using the symmetrized graph.
Chun Hei Michael Chan, Alexandre Cionca, D. Ville· 0 citations
The results show that smaller clusters are generally more vulnerable to attacks on central nodes, whereas larger and less centralized clusters retain more topological efficiency.
The intensity function, defined as the Lebesgue density of the expected measure of a persistence diagram, is a fundamental summary of the probability distribution of persistence diagrams in topological data analysis (TDA). Although several methods have been proposed for estimating intensity functions, statistical hypothesis testing for intensity functions remains largely unexplored. In particular, little is known about the power properties of hypothesis tests based on persistence diagrams. We propose a kernel-based permutation test and analyze its power against alternatives characterized by differences in persistence intensity functions. We introduce assumptions that control the effect of the possibly unbounded cardinality of persistence diagrams and yield a sharp variance bound for the test statistic. We also show that our probability model is broad enough to include all probability densities on the subset of $\mathbb{R}^2$ where $y>x\geq 0$. Using these results, we establish minimax optimality of the proposed test. Along the way, we derive an explicit characterization of the persistence diagram of the \v{C}ech complex on the circle. Since the optimal bandwidth is not directly accessible in practice, we adopt a bandwidth aggregation framework. Simulations and real-data applications demonstrate validity and high empirical power.