To save content items to your account,
please confirm that you agree to abide by our usage policies.
If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account.
Find out more about saving content to .
To save content items to your Kindle, first ensure no-reply@cambridge.org
is added to your Approved Personal Document E-mail List under your Personal Document Settings
on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part
of your Kindle email address below.
Find out more about saving to your Kindle.
Note you can select to save to either the @free.kindle.com or @kindle.com variations.
‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi.
‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.
Problems of measurement error pervade all econometrics. In microeconometrics, a common source of the measurement error problem comes from incorrect response to a survey question, incorrect coding of a correct response, and the use of a correctly measured variable as a proxy for another theoretically valid but unobserved variable (e.g., using observed income as a proxy for “normal income”). Questions that seek sensitive information may elicit partial or incorrect responses. That is, a measurement error is triggered by unobservables (or latent variables) when such variables are replaced by proxy variables.
Here are some examples. Consider the problem of testing for the presence of gender bias in a study of earnings. The obvious approach is to regress a measure of earnings on a categorical gender variable while controlling for qualifications, age, experience, and so forth. However, the most relevant variable may be an individual's on-the-job productivity, which may not be directly observed and a proxy may be used instead. Therefore, the impact of measurement error on inferences about the gender discrimination is an important issue. Studies of individual demand for goods and services feature concepts such as “economic cost” or “full price of a service.” However, these are rarely directly measured in published data and must be constructed by the econometrician prior to model estimation. Inevitably their measurement is subject to error.
There are virtually no models discussed in this book that are protected from the problem of measurement errors.
Exact finite-sample results are unavailable for most microeconometrics estimators and related test statistics. The statistical inference methods presented in preceding chapters rely on asymptotic theory that usually leads to limit normal and chi-square distributions.
An alternative approximation is provided by the bootstrap, due to Efron (1979, 1982). This approximates the distribution of a statistic by a Monte Carlo simulation, with sampling done from the empirical distribution or the fitted distribution of the observed data. The additional computation required is usually feasible given advances in computing power. Like conventional methods, however, bootstrap methods rely on asymptotic theory and are only exact in infinitely large samples.
The wide range of bootstrap methods can be classified into two broad approaches. First, the simplest bootstrap methods can permit statistical inference when conventional methods such as standard error computation are difficult to implement. Second, more complicated bootstraps can have the additional advantage of providing asymptotic refinements that can lead to a better approximation in-finite samples.
Applied researchers are most often interested in the first aspect of the bootstrap. Theoreticians emphasize the second, especially in settings where the usual asymptotic methods work poorly in finite samples.
The econometrics literature focuses on use of the bootstrap in hypothesis testing, which relies on approximation of probabilities in the tails of the distributions of statistics. Other applications are to confidence intervals, estimation of standard errors, and bias reduction.
Cross-section models have certain inherent limitations. They are predominantly equilibrium models that generally do not shed light on intertemporal dependence of events. They also cannot satisfactorily resolve fundamental issues about the sources of persistence in behavior. Such persistence may be behavioral, i.e. arising from true state dependence, or it may be spurious, being an artifact of the inability to control for heterogeneous behavior in the population. Because panel data, also called longitudinal data, contain periodically repeated observations of the same subjects, they have a large potential for resolving issues that cross-section models cannot satisfactorily handle. Chapters 21 through 23 present methods for panel data. We progress systematically from linear models for continuous data in Chapter 21 to nonlinear panel data models for limited dependent variables in Chapter 23. Both fixed effects and random effects models are considered. A persistent theme through these three chapters is the importance of using panel-robust methods of inference.
Chapter 21, which reviews the key general results for linear panel data regression models, can be read easily by those with a good grasp of linear regression; it does not require the material covered in Parts 2 to 4. We recommend that even those who are interested in more advanced material should quickly peruse through the contents of this chapter first to gain familiarity with key concepts and definitions.
Chapter 22 covers important extensions of Chapter 21, especially to dynamic panels which allow for Markovian dependence structure of current variables.
Econometric models of durations are models of the length of time spent in a given state before transition to another state, such as duration unemployed or alive or without health insurance. In biostatistics a duration in a state is also known as lifetime and the time of transition is referred to as death; in operations research where one often studies lifetimes of physical objects such as light bulbs and machines, the end of useful life, that is, transition to useless life, is called failure time. In econometrics a state is a classification of an individual entity at a point in time, transition is movement from one state to another, and a spell length or duration is the time spent in a given state. A typical regression example is determining the effect of higher unemployment benefit levels on the average length of an unemployment spell or the probability of transition out of unemployment.
The literature on this subject can be quite daunting, for a number of reasons. First, several related distributional functions are of interest and either the duration or probability of transition may be modeled. Second, many different sampling schemes are possible and statistical inference depends on both the duration model and the sampling scheme. For example, sampling methods for data on unemployment duration include flow sampling of those entering unemployment in a given month, stock sampling of people unemployed in a given month, and population sampling of all people regardless of employment status.
Panel data are repeated observations on the same cross section, typically of individuals or firms in microeconomics applications, observed for several time periods. Other terms used for such data include longitudinal data and repeated measures. The focus is on data from a short panel, meaning a large cross section of individuals observed for a few time periods, rather than a long panel such as a small cross section of countries observed for many time periods.
A major advantage of panel data is increased precision in estimation. This is the result of an increase in the number of observations owing to combining or pooling several time periods of data for each individual. However, for valid statistical inference one needs to control for likely correlation of regression model errors over time for a given individual. In particular, the usual formula for OLS standard errors in a pooled OLS regression typically overstates the precision gains, leading to underestimated standard errors and t-statistics that can be greatly inflated.
A second attraction of panel data is the possibility of consistent estimation of the fixed effects model, which allows for unobserved individual heterogeneity that may be correlated with regressors. Such unobserved heterogeneity leads to omitted variables bias that could in principle be corrected by instrumental variables methods using only a single cross section, but in practice it can be difficult to obtain a valid instrument.
A great deal of empirical microeconometrics research uses linear regression and its various extensions. Before moving to nonlinear models, the emphasis of this book, we provide a summary of some important results for the single-equation linear regression model with cross-section data. Several different estimators in the linear regression model are presented.
Ordinary least-squares (OLS) estimation is especially popular. For typical microeconometric cross-section data the model error terms are likely to be heteroskedastic. Then statistical inference should be robust to heteroskedastic errors and efficiency gains are possible by use of weighted rather than ordinary least squares.
The OLS estimator minimizes the sum of squared residuals. One alternative is to minimize the sum of the absolute value of residuals, leading to the least absolute deviations estimator. This estimator is also presented, along with extension to quantile regression.
Various model misspecifications can lead to inconsistency of least-squares estimators. In such cases inference about economically interesting parameters may require more advanced procedures and these are pursued at considerable length and depth elsewhere in the book. One commonly used procedure is instrumental variables regression. The current chapter provides an introductory treatment of this important method and additionally addresses the complication of weak instruments.
Section 4.2 provides a definition of regression and presents various loss functions that lead to different estimators for the regression function. An example is introduced in Section 4.3.
A nonlinear estimator is one that is a nonlinear function of the dependent variable. Most estimators used in microeconometrics, aside from the OLS and IV estimators in the linear regression model presented in Chapter 4, are nonlinear estimators. Nonlinearity can arise in many ways. The conditional mean may be nonlinear in parameters. The loss function may lead to a nonlinear estimator even if the conditional mean is linear in parameters. Censoring and truncation also lead to nonlinear estimators even if the original model has conditional mean that is linear in parameters.
Here we present the essential statistical inference results for nonlinear estimation. Very limited small-sample results are available for nonlinear estimators. Statistical inference is instead based on asymptotic theory that is applicable for large samples. The estimators commonly used in microeconometrics are consistent and asymptotically normal.
The asymptotic theory entails two major departures from the treatment of the linear regression model given in an introductory graduate course. First, alternative methods of proof are needed since there is no direct formula for most nonlinear estimators. Second, the asymptotic distribution is generally obtained under the weakest distributional assumptions possible. This departure was introduced in Section 4.4 to permit heteroskedasticity-robust inference for the OLS estimator. Under such weaker assumptions the default standard errors reported by a simple regression program are invalid. Some care is needed, however, as these weaker assumptions can lead to inconsistency of the estimator itself, a much more fundamental problem.
In this chapter we consider two closely related topics: regression when the dependent variable of interest is incompletely observed and regression when the dependent variable is completely observed but is observed in a selected sample that is not representative of the population. This includes limited dependent variable models, latent variable models, generalized Tobit models, and selection models.
All these models share the common feature that even in the simplest case of population conditional mean linear in regressors, OLS regression leads to inconsistent parameter estimates because the sample is not representative of the population. Alternative estimation procedures, most relying on strong distributional assumptions, are necessary to ensure consistent parameter estimation.
Leading causes of incompletely observed data are truncation and censoring. For truncated data some observations on both the dependent variable and regressors are lost. For example, income may be the dependent variable and only low-income people are included in the sample. For censored data information on the dependent variable is lost, but not data on the regressors. For example, people of all income levels may be included in the sample, but for confidentiality reasons the income of high-income people may be top-coded and reported only as exceeding, say, $100,000 per year. Truncation entails greater information loss than does censoring. A leading example of truncation and censoring is the Tobit model, named after Tobin (1958), who considered linear regression under normality.
This book provides a detailed treatment of microeconometric analysis, the analysis of individual-level data on the economic behavior of individuals or firms. A broader definition would also include grouped data. Usually regression methods are applied to cross-section or panel data.
Analysis of individual data has a long history. Ernst Engel (1857) was among the earliest quantitative investigators of household budgets. Allen and Bowley (1935), Houthakker (1957), and Prais and Houthakker (1955) made important contributions following the same research and modeling tradition. Other landmark studies that were also influential in stimulating the development of microeconometrics, even though they did not always use individual-level information, include those by Marschak and Andrews (1944) in production theory and by Wold and Jureen (1953), Stone (1953), and Tobin (1958) in consumer demand.
As important as the above earlier cited work is on household budgets and demand analysis, the material covered in this book has stronger connections with the work on discrete choice analysis and censored and truncated variable models that saw their first serious econometric applications in the work of McFadden (1973, 1984) and Heckman (1974, 1979), respectively. These works involved a major departure from the overwhelming reliance on linear models that characterized earlier work. Subsequently, they have led to significant methodological innovations in econometrics. Among the earlier textbook-level treatments of this material (and more) are the works of Maddala (1983) and Amemiya (1985).
Part 1 emphasized that microeconometric models are frequently nonlinear models estimated using large and heterogeneous data sets drawn from surveys that are complex and subject to a variety of sampling biases. A realistic depiction of the economic phenomena in such settings often requires the use of models for which estimation and subsequent statistical inference are difficult. Advances in computing hardware and software now make it feasible to tackle such tasks. Part 3 presents modern, computerintensive, simulation-based methods of estimation and inference that mitigate some of these difficulties. The background required to cover this material varies somewhat with the chapter, but the essential base is least squares and maximum likelihood estimation.
Chapter 11 presents bootstrap methods for statistical inference. These methods have the attraction of providing a simple way to obtain standard errors when the formulae from asymptotic theory are complex, as is the case, for example, for some two-step estimators. Furthermore, if implemented appropriately, a bootstrap can lead to a more refined asymptotic theory that may then lead to better statistical inference in small samples.
Chapter 12 presents simulation-based estimation methods. These methods permit estimation in situations where standard computational methods may not permit calculation of an estimator, because of the presence of an integral over a probability distribution that leads to no closed-form solution.
Chapter 13 surveys Bayesian methods that provide an approach to estimation and inference that is quite different from the classical approach used in other chapters of this book.