Skip to content
Review Open access

BUILDING A CLUSTER LINEAR REGRESSION WITH CONSTRAINTS ON THE SIZE OF SUBSAMPLES OF DATA

2026 · Informatika i sistemy upravleniya · 0 citations

Abstract

The paper provides a brief overview of publications related to the division of the initial data sample into disjoint fragments when modeling complex objects. In particular, there is considered distributed asymmetric least squares estimation based on a Poisson subsample; a new methodology for the structural synthesis of specialized parallel computational sub-blocks for implementing a group method of data processing algorithms; ANOVA class method based on a dedicated subsample of data, which suggests ways to reduce the effect of biased variance estimation; a model of the intensity of occurrence and size distribution of forest fires, combining the theory of extreme values and point processes within the framework of a new Bayesian hierarchical model; recent advances in theory, methods and implementations of quantile regression in the context of massive and streaming data; methods for reducing the size of the initial sample due to missing values. The authors formulate the task of identifying the parameters of a cluster regression model with predefined capacities of each allocated subsample of the initial data sample. When assigning the sum of the error modules of the approximation as a loss function for calculating the model parameters, this task is reduced to a linear Boolean programming problem. The case is considered when the sizes of the formed clusters are approximately equal. Two new variants of the cluster regression model for the development of the chemical industry in the Russian Federation have been constructed.

Read PDF