Missing values, atypical observations, and heterogeneity across latent groups are common sources of complexity in regression data. The contaminated Gaussian cluster-weighted model (CG-CWM) provides a natural framework for handling atypical observations, including outliers and leverage points, in model-based clustering. We extend the CG-CWM to data with missing-at-random (MAR) values in both the response and covariate spaces. The proposed model provides clustering in regression analysis while distinguishing typical observations, outliers, and good and bad leverage points. By treating covariates as random, the model preserves assignment dependence, allowing them to contribute directly to cluster formation. Maximum likelihood estimation is performed through an expectation-conditional maximization (ECM) algorithm that accounts for four sources of incomplete information: missing responses and covariates, unknown component memberships, and latent contamination indicators. Conditional on these indicators, the joint distribution of responses and covariates is multivariate Gaussian, yielding closed-form conditional distributions for missing values and incorporating missingness uncertainty directly into parameter updates. Thus, missing values are handled within model fitting rather than by preliminary imputation. The framework provides clustering, clusterwise regression, model-based treatment of MAR values, and detection of atypical observations. Performance is assessed through numerical studies under varying levels of contamination and missingness patterns, and a real data application.
Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data and not on unobserved data. This paper addresses the problem of estimating a multivariate linear regression model with multiple random covariates in the presence of MAR values in both the response and covariate spaces using a maximum likelihood (ML) framework. The proposed methodology models the joint distribution of responses and covariates through a conditional-marginal factorization of a multivariate Gaussian distribution. This formulation can be interpreted as a reparameterization of the multivariate normal distribution when the variables can be naturally partitioned into responses and covariates. Parameter estimation is performed using the expectation-maximization (EM) algorithm, which facilitates the imputation of missing values while preserving the distinct roles of responses and covariates. We extend this framework to the model-based clustering setting by considering a mixture of multivariate linear regressions with multiple random covariates. This extension enables soft clustering under incomplete data and accommodates MAR values in both the multivariate responses and covariates. Hence, it represents one of the most general model-based clustering solutions for regression data currently available in the literature. The effectiveness of the methodology is demonstrated through a simulation study, and the advantages of the proposed reparameterization are illustrated using the Automobile dataset, which contains missing values.
Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.
J. Pillay, A. Bekker, C. Tortora et al.· 1 citation
Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.
J. Pillay, A. Bekker, C. Tortora et al.· 0 citations