Frisch–Waugh–Lovell theorem

Theorem in statistics and econometrics From Wikipedia, the free encyclopedia

In statistics and econometrics, the Frisch–Waugh–Lovell[a] (FWL) theorem proves a property of ordinary least squares estimators. The theorem is named for econometricians Ragnar Frisch, Frederick V. Waugh, and Michael C. Lovell.

Ragnar Frisch, econometrician and founding member of the Econometric Society

Ordinary least squares is a method of estimating coefficients in a linear regression, where a single dependent variable is modeled as a linear function of one or more independent variables. The Frisch–Waugh–Lovell theorem states that, in a least squares-estimated regression, each independent variable's coefficient reflects the relationship between the dependent variable and the part of that independent variable which is not linearly related to the other independent variables. Specifically, each independent variable can be decomposed into two parts: the part linearly related to the other independent variables, and a residual component. Then, that variable's coefficient can be found by regressing the dependent variable on the residual component.

Coefficients in least squares-estimated regressions are often interpreted as the effect of the respective variable controlling for or holding constant the set of other independent variables. The Frisch–Waugh–Lovell theorem formalizes this, showing what least squares does in controlling for variables. As a result, the theorem is sometimes called the regression anatomy theorem.

An initial version of the theorem was introduced by Udny Yule in 1907, though it was not popularized in economics until a 1933 paper by Ragnar Frisch and Frederick Waugh in the first volume of Econometrica. At the time, there was debate among economists over the proper way to adjust statistical models for the effect of time trends. Frisch and Waugh used the theorem to show that the two leading methods of adjustment were numerically equivalent, resolving the debate. Michael Lovell contributed to the theorem's development through a 1963 paper generalizing Frisch and Waugh's result beyond time trends to arbitrary sets of independent variables.

Background

The Frisch–Waugh–Lovell theorem is a result for regressions estimated by ordinary least squares, the most commonly used estimator in applied econometrics.[1] A regression is a statistical model where a dependent variable is modeled as a function of one or more independent variables plus some residual term. When a regression is linear in parameters – that is, each variable's contribution to the predicted value is linear – ordinary least squares can be used to estimate the model's coefficients. More formally, ordinary least squares can be used when a dependent variable is modeled as a linear combination of one or more independent variables plus some residual term. For example, an individual's wages may be modeled as a linear function of a constant term, education, and parental income, with a residual term that encompasses deviations from the model's prediction. Ordinary least squares sets the coefficients to minimize the sum of squared residuals.[2] Under a certain set of assumptions, the hypotheses of the Gauss–Markov theorem, least squares estimation is the best linear unbiased estimator.[3]

Let be any dependent variable and a set of independent variables, and suppose observations of are obtained. If is modeled as a linear function of the independent variables and a constant, the estimated model can be written as , where the hats denote estimates of the respective parameters. The least squares estimator sets the coefficients to minimize the sum of squared residuals. With observations this involves minimizing across equations, and is typically written in matrix form as , where and are vectors of dependent variable observations and residuals, respectively, is an matrix of independent variables' observations, and is a coefficient vector. Then, the least squares solution is yielding .[4]

In regressions estimated by least squares, it is common to refer to an independent variable's coefficient as the effect of that variable "holding constant" the other independent variables.[5] For example, if wage is modeled as a function of education and work experience, the coefficient on education is interpreted as the difference in the expectation of wage for a unit difference in education, "holding constant" work experience. Econometrician Arthur Goldberger frames the Frisch–Waugh–Lovell theorem as "giving content to th[is] language".[6]

Definition and interpretation

The Frisch–Waugh–Lovell theorem states that in a least squares-estimated regression of the form

any coefficient can be obtained by the two-step process of:

  1. Regress on the set of other independent variables, obtaining residuals
  2. Regress on , obtaining

This two-step process is referred to as the residual regression or equivalently the regression anatomy theorem.[7][8] This result is a numerical property of least squares estimation and does not depend on statistical properties of the data.[9][10]

From this theorem, each independent variable in a least squares-estimated regression can be decomposed into two parts: the part which is linearly related to the set of other independent variables, and the residual which remains. Then, that independent variable's coefficient can be found from the simple regression of the dependent variable on the residual part.[11][12] This result is the basis for interpreting the impact of including additional variables in a regression: it is equivalent to removing from the existing variables the part which the new variables linearly explain.[13][14]

The Frisch–Waugh–Lovell theorem can, for example, be applied to interpret multicollinearity. When most of the variation in an independent variable is linearly explained by the other independent variables, the residual part has relatively very little variation. Resultingly, the estimate of the independent variable's coefficient may be less precise than if fewer variables were controlled for.[15][16]

Example

Consider the regression of wage on education and parental income:

While the least squares estimates for and can be obtained by minimizing directly, each can be equivalently obtained by the two-step residual regression. In the case of education:

  1. Regress education on parental income, saving the residuals from this regression: the part of education not linearly related to parental income
  2. Regress wages on the residuals, obtaining the least squares estimate for

This illustrates how reflects the effect of education on wages controlling for parental income: it is the relationship between wages and the part of education not linearly related to parental income.[8]

Double residual regression

The double residual regression is the three-step process:

  1. Regress on the set of other independent variables, obtaining residuals
  2. Regress on the set of independent variables excluding , obtaining residuals
  3. Regress on , estimating and

Like the two-step process, this yields an identical coefficient to the full regression.[17][18] It includes the additional feature that the residuals from the regression in step 3 equal the residuals in the full regression.[11][19]

Multivariate definition

Consider the least squares-estimated regression , where and are vectors of dependent variable observations and residuals, respectively, is an matrix of independent variables' observations, is an matrix of independent variables' observations, and and are and coefficient vectors for and , respectively. Then, the Frisch–Waugh–Lovell theorem states that

where , the residuals from the least squares regression of on , and , the residuals from the least squares regression of on . The first expression of is the residual regression, while the second is the double residual regression.[20][21][22]

Geometric interpretation

With a linear regression of the form , the fitted values can be interpreted as the orthogonal projection of onto the column space of , .[23] The Frisch–Waugh–Lovell theorem is then (in the double residual regression case) the three step process:

  1. Project onto the orthogonal complement of , obtaining residuals
  2. Project onto the orthogonal complement of , obtaining residual vector
  3. Project onto , obtaining projection and residuals

The resulting and residuals are identical to those in the full regression of on and .[24][25][26]

Proof

Consider the least squares-estimated regression and annihilator matrix . Premultiplying both sides of the regression equation by the annihilator matrix removes from and the part linearly explained by :

Where the third line follows from by construction and from , a property of least squares. Then, by the least squares result, and . This concludes the proof.[6][27][28]

History

Yule

In 1907, statistician Udny Yule introduced a new system of notation for, and derived a number of algebraic results of, least squares-estimated regression coefficients. Among his results was an early form of the Frisch–Waugh–Lovell theorem.[29][30] In Yule's notation, where represents the residuals from the regression of on through , and the residuals from the regression of on through , he finds that the regression of on yields the coefficient for in the full regression of on through . He notes that this relationship holds "quite generally and without reference to the form of the [variables'] frequency distribution." [31] With this result, Yule defines the multiple regression coefficient – the coefficient on in the regression of on through – as the simple regression of on :

Having related simple regression coefficients to multiple regression coefficients, Yule describes his result as filling a gap in the interpretation of least squares coefficients and partial correlations by showing that they reflect "an actual correlation between determinate variables."[31][32]

Using Yule's notation, in a 1968 text econometrician Arthur Goldberger states the residual regression form of the Frisch–Waugh–Lovell theorem as and the double residual regression form as .[33]

Frisch and Waugh

In the early 20th century, there was debate among economists over the correct approach to adjusting time series data used in regressions for the influence of linear trends. The two primary methods in question were the direct de-trending of each time series and the inclusion of a time trend in the regression. In 1933 and using the notation introduced by Yule, a paper in the first volume of Econometrica by econometricians Ragnar Frisch and Frederick V. Waugh proved the equivalence between the two methods.[34][35][36]

Prior to Frisch and Waugh's result, much of the debate around the optimal time trend adjustment concerned estimates of static demand equations whose observations had been taken over time. Because the theoretical demand equations did not depend on time, economists sought to adjust for the effect of time trends in order to bring the statistical model closer in line with theory. Advocates of including time trends in regressions argued it improved the model's fit, where opponents argued that time trends may violate ceteris paribus assumptions of the underlying theoretical model.[37] In proving the equivalence of the two methods, addressing the difference in model fit, and formalizing the distinction between estimated coefficients and theoretical models, Frisch and Waugh's paper resolved the debate around trend adjustments.[38]

Economist and historian Mary S. Morgan contextualizes Frisch and Waugh's result, as it pertains to a greater understanding of regression coefficients, as having "paved the way for a more generous use of the other factors in the demand equation."[39] Frisch and Waugh's results were, in 1952, extended by Gerhard Tintner to polynomial trend adjustment.[40][41] In a 1953 textbook on demand analysis, econometrician Herman Wold references Frisch and Waugh's paper as a special case applied to time adjustments.[42]

Generalization and later development

In 1963, econometrician Michael C. Lovell generalized Frisch and Waugh's results, providing a proof of the theorem in matrix notation and applying it to seasonality adjustments.[29][24][43] Rather than focusing on certain types of variables, as Frisch and Waugh did with time trends, Lovell proves the result with arbitrary sets of independent variables.[44] Lovell presents 7 regression specifications and proves how their coefficients relate, among them both the residual and double residual forms of the theorem.[22] Lovell published an additional proof in 2008 using only simple algebra.[44]

In 1964, economist Richard Stone published a generalized proof of the theorem.[45]

The Frisch–Waugh–Lovell theorem is included in most intermediate to advanced econometrics textbooks.[24]

Naming

The theorem has been referred to under a number of names, including the Frisch–Waugh–Lovell theorem, Frisch-Waugh theorem, partitioned regression theorem, residual regression, and the regression anatomy theorem.[24][15]

While Frisch and Waugh's paper was not the first introduction of the result, it was the first proof in econometrics.[32][24] Recognizing the generalization by Lovell, the theorem was presented as the Frisch–Waugh–Lovell theorem in a 1993 econometrics textbook by Russell Davidson and James G. MacKinnon.[24]

Extensions

Where the Frisch–Waugh–Lovell theorem states that the full and residual regressions have the same coefficients, relationships between the coefficients' standard errors can also be shown. Lovell's 1963 paper finds that the homoskedastic standard errors of coefficients in the double residual regression differ from those of the full regression by a degrees of freedom adjustment.[46] In 2021, statistician Peng Ding presented a proof of Lovell's results and found comparable results for other estimates of standard errors, including heteroskedasticity-consistent and clustered standard errors.[47]

Analogues to the Frisch–Waugh–Lovell theorem have been shown for a number of other estimators, including generalized least squares,[48] ridge regression and the LASSO,[49] and k-class estimators, including limited information maximum likelihood.[50]

See also

Notes

References

Related Articles

Wikiwand AI