Bayesian DAG Structure Learning with Simultaneous Shrinkage Covariance Estimation under Scale-Mixture Error Distributions in the Proportional High-Dimensional Regime
We propose a unified Bayesian framework namely robust DAG-Cholesky horseshoe (R-DACH) for joint directed acyclic graph (DAG) structure learning and precision matrix estimation in the high-dimensional proportional asymptotic regime $p/n \to c \in (0,\infty)$, under the scale mixture of normal errors. The construction places a global-local horseshoe-type prior directly on the strictly lower-triangular entries of the modified Cholesky factor of the DAG-Markov precision matrix, so that sparsity in the Cholesky parameters induces a coherent parent-set selection consistent with a topological ordering of the variables. A per-observation inverse-gamma scale mixture yields automatic robustness to heavy-tailed and contaminated observations and admits Student-$t$, Laplace, and slash distributions as special cases. We design a partially-collapsed blocked Gibbs sampler that traverses the joint space of orderings, sparsity patterns and continuous parameters. Simulations across $(n,p)$ configurations with $p$ up to several hundreds confirm the theoretical rates and demonstrate substantial gains over graphical-horseshoe, DAG-Wishart, and PC-based competitors under contamination. An application to RNA-seq gene-expression data from \emph{The Cancer Genome Atlas} reveals biologically interpretable regulatory structure that competing methods fail to recover.
This work proposes SURE-Ridge, a non-iterative, closed-form estimator for equal variance linear Gaussian SEM, which achieves the lowest structural Hamming distance in the small-sample regime and the lowest run time across all sample sizes tested, compared with NOTEARS, DAGMA, and GBNSL baselines.
Structure learning of directed acyclic graphs (DAGs) from observational data is a foundational task in causal discovery and is widely used to infer regulatory networks from medical and genomic measurements. The Bayesian formulation quantifies model uncertainty and admits prior biological knowledge, but its practical use has been hampered by the super-exponential growth of the DAG space and by the intractability of the node-marginal likelihood under flexible, non-conjugate priors. Existing closed-form solutions are largely confined to the conjugate Normal--Inverse-Gamma prior. We develop a Laplace-approximated Bayesian scoring function for the non-conjugate Normal--Gamma prior on the modified Cholesky parameterisation of the precision matrix, embed it in a Metropolis--Hastings sampler over DAGs, and couple the latent Gaussian network to a binary clinical outcome through a probit link. We show that the node-marginal integral is of generalised inverse-Gaussian form, so that its exact value is a modified Bessel function of the second kind and the proposed scoring function is its leading large-argument asymptotic; the posterior of each conditional variance is likewise generalised inverse-Gaussian and is sampled exactly. In simulation, the proposed prior improves on the conjugate baseline and on the PC, greedy-equivalence-search, NOTEARS, and DAGMA benchmarks at sample sizes typical of clinical cohorts. On two real datasets, the Sachs protein-signalling network, scored against its validated consensus graph, and the Wisconsin Diagnostic Breast Cancer data, the method recovers known structure and, through the DAG-probit extension, predicts malignancy from nuclear morphometry with a cross-validated ROC-AUC of $0.94$ using a sparse, interpretable set of direct predictors.
Samaneh Nazari, Mohammad Arashi, Abdolnasser Sadeghkhani· 0 citations
Hierarchical mixture models are a powerful tool for modeling data generated from heterogeneous sources, particularly when the mixing proportion $\boldsymbol{w}$ itself is treated as a random variable with a Dirichlet or Beta-Liouville prior. Such models are widely employed in scenarios where uncertainty in class membership or data-generating processes must be probabilistically quantified. This paper studies the exact marginalization of the mixture weight. For the two-component case we give an $O(n^2)$ dynamic program -- and an $O(n \log^2 n)$ FFT variant -- for the marginal likelihood, and show that the exact posterior of the weight is a finite mixture of Beta distributions, delivering closed-form posterior summaries, credible intervals and per-observation local false-discovery rates without any sampling. For $K \ge 3$ components we give an exact joint dynamic program. The gain is largest in the small-sample regime the method is built for: on a real multilevel meta-analysis, a pathway-level dysregulation analysis of leukemia gene expression, and a leukemia-derived gene-panel benchmark with known ground truth, the exact interval for the signal proportion is calibrated where EM gives no interval at all (collapsing to a boundary) and Gaussian/Laplace approximations mis-cover, and it is two orders of magnitude faster than the sampler that would match it. On the large prostate-cancer benchmark, where every method has ample data, it agrees with locfdr on the gene ranking while adding a posterior interval for the null proportion.
We develop a computationally scalable Bayesian framework for precision matrix estimation in Gaussian graphical models under total positivity constraints. To overcome the high computational cost of the Gaussian likelihood, we adopt a generalized Bayesian approach based on the $D$-trace loss, which eliminates the log-determinant term and enables efficient optimization while allowing relaxation of positive definiteness during sampling. Sparsity is induced via spike-and-slab priors, and the resulting generalized posterior is shown to be proper under mild conditions. Our primary contribution is a suite of efficient posterior sampling algorithms tailored to high-dimensional settings. Starting from a component-wise Gibbs sampler, we introduce a novel data augmentation scheme that induces conditional independence among precision matrix entries, enabling joint updates. By exploiting the Gram structure of the sample covariance matrix, we further develop a fast matrix-normal sampler that significantly reduces per-iteration complexity in high-dimensional settings. An interweaving strategy combines augmented and direct updates to improve mixing without sacrificing scalability. Experiments on synthetic and financial data demonstrate substantial computational gains over existing methods, while maintaining competitive estimation accuracy and improved recovery of structured dependencies.
Swarnali Raha, Partha Sarkar, Sirani M. Perera et al.· 0 citations
Bayesian quantile regression based on the asymmetric Laplace (AL) distribution can be sensitive to extreme observations because of its exponentially decaying tails. We propose a robust error distribution constructed as a finite mixture of the AL distribution and a log-Pareto scale mixture of asymmetric Laplace distributions (LPAL). Unlike a direct log-Pareto extension of the normal location-scale representation of the AL distribution, the proposed AL-LPAL mixture preserves the prescribed quantile and exhibits log-regularly varying behavior in both tails. The LPAL component also has an unbounded density at the target quantile, yielding a distribution that combines sharp central concentration with super-heavy tails. We establish posterior robustness under arbitrarily extreme contamination and provide sufficient conditions for the existence of posterior moments of the regression coefficients and scale parameter. For posterior computation, we develop a Gibbs sampler using latent-variable augmentations and a computationally efficient mean-field variational Bayes approximation. Simulation studies show that the proposed method is competitive under moderate contamination and maintains stable point estimation with comparatively concentrated posterior intervals, particularly when severe contamination affects the quantile of interest. Applications to carbon dioxide and Boston housing data, using the same preprocessing as existing robust Bayesian quantile regression analyses, show favorable predictive performance across nearly all quantile levels and loss criteria considered.
Dongu Han, Genya Kobayashi, S. Sugasawa· 0 citations
A new approach is introduced that estimates a global structure while accounting for local cluster-level effects, and presents a differentiable graph coupling mechanism that guarantees the union of the fixed- and random-effects graphs remains acyclic.
Ryan Thompson, Matt P. Wand, V. Baladandayuthapani· 0 citations