Home /permanent

Variable Subset Selection

Variable Subset Selection is a technique for estimating parameters of a Linear Model, where we identify a subset of independent variables that are most predictive of the dependent variable.

Consider the following linear model:

y=β0+β1x1+β2x2+…+βpxp+ϵ y = \beta_0 + \beta_1x_1 + \beta_2x_2 + \ldots + \beta_px_p + \epsilon

In this context, yy represents the dependent variable, β0\beta_0 is the intercept, β1,β2,…,βp\beta_1, \beta_2, \ldots, \beta_p are the coefficients, x1,x2,…,xpx_1, x_2, \ldots, x_p are the independent (predictor) variables, and ϵ\epsilon denotes the error term.

Variable subset selection aims to identify a subset of the pp predictor variables x1,x2,…,xpx_1, x_2, \ldots, x_p of size dd that are most strongly associated with the dependent variable. By narrowing down the predictors, we simplify the model while retaining its predictive power.

Once we determine this smaller subset of predictors, we fit a least squares linear regression model using only these variables. For example, if we believe that only x1x_1 and x2x_2 are significantly related to yy, we set d=2d = 2 and fit a model as follows:

y=β0+β1x1+β2x2+ϵ y = \beta_0 + \beta_1x_1 + \beta_2x_2 + \epsilon

This approach assumes that the remaining variables x3,…,xpx_3, \ldots, x_p do not significantly contribute to explaining the variability in the response yy.

But how do we determine which variables are important? In some cases, prior data analysis or domain expertise can guide the selection process. More commonly, however, a systematic approach is needed to assess the relevance of each variable.

The number of potential models that can be formed depends on the total number of predictors. Including the intercept term β0\beta_0, each of the pp predictors x1,…,xpx_1, \ldots, x_p can either be included in or excluded from the model, resulting in 2p2^p possible combinations to consider.