Estimation and Inference for Multi-Threshold Regression with Endogeneity
Heterogeneity and endogeneity become increasingly common in statistical modeling and econometric practice. Threshold regression provides a simple yet flexible modeling strategy to account for heterogeneity. This paper studies estimation and inference for multi-threshold regression with endogenous regressors. We exploit a novel two-stage estimation procedure based on instrumental variables, which first identifies a set of possible threshold locations using the group LASSO estimation, and then refines the candidate set by a predetermined selection criterion. Given that the performance of the conventional information criteria is sensitive to the choice of the penalization factor, we develop a data-adaptive threshold-based cross-validation criterion incorporating an order-preserved sample-splitting strategy to determine the number of thresholds. Regarding inference, we develop test statistics to test for the presence of threshold effects and the existence of endogeneity, respectively. Numerical simulations and an application analyzing the threshold effects of 401(k) plans on wealth demonstrate the excellent finite sample performance of our methods. Finally, we develop a user-friendly R package MultiThreshold to implement the methodologies.
- Research Article
13
- 10.2139/ssrn.2533013
- Dec 3, 2014
- SSRN Electronic Journal
Threshold Regression with Endogeneity
- Research Article
6
- 10.1017/s0266466623000014
- Mar 16, 2023
- Econometric Theory
This paper studies control function (CF) approaches in endogenous threshold regression where the threshold variable is allowed to be endogenous. We first use a simple example to show that the structural threshold regression (STR) estimator of the threshold point in Kourtellos, Stengos and Tan (2016, Econometric Theory 32, 827–860) is inconsistent unless the endogeneity level of the threshold variable is low compared to the threshold effect. We correct the CF in the STR estimator to generate our first CF estimator using a method that extends the two-stage least squares procedure in Caner and Hansen (2004, Econometric Theory 20, 813–843). We develop our second CF estimator which can be treated as an extension of the classical CF approach in endogenous linear regression. Both these approaches embody threshold effect information in the conditional variance beyond that in the conditional mean. Given the threshold point estimates, we propose new estimates for the slope parameters. The first is a by-product of the CF approach, and the second type employs generalized method of moment (GMM) procedures based on two new sets of moment conditions. Simulation studies, in conjunction with the limit theory, show that our second CF estimator and confidence interval for the threshold point together with the associated second GMM estimator and confidence interval for the slope parameter dominate the other methods. We further apply the new estimation methodology to an empirical application from international trade to illustrate its usefulness in practice.
- Research Article
12
- 10.1097/ede.0b013e3181a0acc7
- May 1, 2009
- Epidemiology
Missing covariate data is a feature common to most epidemiologic and clinical studies. It is now widely recognized that the routinely used complete-case analysis—which excludes subjects with any missing variable—can (i) yield highly inefficient estimators in regression analysis, and (ii) be severely biased when the data are not a completely random sample of the full data (thus not missing completely at random). In the simplest and most familiar setting where the observed data arise from simple random sampling (ie, independent and identically distributed data), weighted estimating equations provide a powerful framework to perform regression analysis, while appropriately accounting for covariate data missing at random but not necessarily missing completely at random. In this issue of Epidemiology, Moore et al1 use 2-stage weighted estimating equations or re-weighted estimating equations to perform logistic regression of Y, the indicator of “trying to loose weight in the last 12 months” on Z, the indicator of high cholesterol that is missing by happenstance, controlling for a subset of fully observed covariates X collected in the third National Health and Nutrition Examination Survey (NHANES III). NHANES III used stratified, multistage sampling and thus does not follow a simple random sampling design. The authors use all sampled participants with complete observations, but propose to jointly account for individuals' differential propensity both to be selected into the survey sample and to have an observed cholesterol indicator. They note that this can be achieved by multiplying participants' respective contribution to the complete-case logistic score equation by the product of the (known) inverse probability of being selected and the (unknown) inverse probability of having an observed cholesterol indicator (see equation (3) of their paper). The missingness mechanism, which stochastically determines which of the sampled participants have complete data, is unknown to the analyst and therefore must be estimated. Fortunately, missingness at random implies that an individual's probability of having an observed cholesterol indicator (R = 1) does not depend on whether or not she has high cholesterol, but may depend on other fully observed variables Y, X, and S. Here, S includes covariates that might not be of primary scientific interest, but are necessary with Y and X, to explain any association between R and Z. Under this assumption, one can therefore proceed by substituting the unknown weights with estimated weights,2 constructed with an estimate ˆπ = π (X, Y, S; ˆα) of Pr(R = 1|X, Y, S), where ˆα denotes the observed data maximum likelihood estimator of the coefficients of the logistic regression of R on (X, Y, S). Moore et al essentially adapt this approach to their two-stage procedure, but require the stronger assumption that S can be dropped from π. In this commentary, special consideration is given to issues of statistical efficiency and modeling robustness. First, even in the absence of model mis-specification, inverse-probability weighting can be highly inefficient. This is because, even when (as suggested by Moore et al), one fits a highly parameterized model for π (of course within limits of the data), inverse-probability weighting still fails to make optimal use of data (Y, X, S) observed among all individuals including those missing Z. Second, an incorrect working regression model for the missingness mechanism, π, will invariably result in a biased re-weighted estimator of the parameter β = (β0, βx′, βz) indexing the main logistic regression model of interest logit(Pr(Y = 1|X, Z; β)) = β0 + βx′X + βzZ. In this short commentary, we briefly address these 2 issues by describing a straightforward implementation of an estimator that (a) partially protects against model mis-specification of the weights by remaining consistent for β when the model for π is incorrect, provided that a working parametric model g (Z|Y, X, S; η), with unknown parameter η, for the conditional density of the covariate Z given (Y, X, S), is correct, (b) remains consistent for β when π is correctly modeled, irrespective of the correctness of the posited model for the conditional density of Z, and (c) is in large sample and in the absence of model mis-specification, guaranteed to be more efficient than inverse-probability weighting. By virtue of satisfying (a) and (b), this estimator is an example of a so-called doubly robust estimator. Because we rarely know which if any of (a) or (b) holds, double robustness is a desirable property as it effectively provides the analyst with 2 separate chances of obtaining a consistent estimator of the logistic regression parameter β, instead of the single chance conferred by inverse-probability weighting and conventional likelihood or Bayesian solutions to this problem.3–8 When compared with existing implementations of doubly robust estimators,8 the main appeal of the proposed approach is that it is easily performed using Proc NLMIXED in SAS and thus can be used in routine data analyses without much additional effort. The accompanying technical report9 provides an extension of the method to regression analysis with outcome missing at random, and to the estimation of a marginal structural mean model for point exposure under the assumption of no unmeasured confounding. A Simple Doubly Robust Estimator To simplify this presentation, I depart from the survey sampling framework of the paper; however, as I will later discuss, the described methodology equally applies to the survey setting. Thus, I assume that the observed data consists of n samples of independent and identically distributed data O = (R, ZR, X, S, Y) obtained from a common statistical distribution function. Then, under the missing at random assumption, and assuming that an individual's probability of having complete data is bounded away from 0, ˆβdr satisfies (a), (b), and (c), where ˆβdr is obtained as followed: First, obtain the estimator ˆη that maximizes the complete-case working log-likelihood function σni=1Ri × log (g (Zi|Yi, Xi, Si; η)). For Zi categorical, continuous or a count, conventional models from the exponential family as well as any other user-supplied parametric model can be implemented in Proc NLMIXED of SAS. If Zi is binary, go to the next step; otherwise, for observation i = 1, … n, and user-supplied integer K, generate pseudo data Zi,k*, k = 1,…, K, from ĝ = g(Zi,k*|Yi, Xi, Si; ˆη). Then go to the next step. For binary Zi, find ˆβdr that maximizes the weighted-pseudo-log-likelihood function ll (β*) = σi=1n[ˆωi log(f (Yi|Xi, Zi; β*)) + (1 − ˆωi) × (σz=01g(z|Yi, Xi, Si; ˆη) × log (f (Yi|Xi,z; β*)))] where log (f (Yi|Xi, z; β*)) = (Yi log (Pr (Y = 1|Xi, Zi; β*)) + (1 − Yi) log (Pr(Y = 0|Xi, Zi; β*)) is the log-likelihood for logistic regression, and Categorical Zi can generally be handled in a fashion similar to the binary case, provided the number of categories is not excessive. For other types of variable Zi, let ˆβdr = ˆβdr (K) be the maximizer of llK (β*) = σi=1n[ˆωif (Yi|Xi, Zi; β*) + (1 − ˆωi) × (σk=1Kf (Yi|Xi, Zi,k*; β*))], where Zi,k* k = 1, . . ., K are obtained in step 2. Maximizing either ll(β*) or llK(β*) can be done in SAS Proc NLMIXED with the “model y∼ general (·)” statement, where · stands for an individual's contribution to either ll (β*) or llK (β*) whose functional form can be specified by the user. See the example of programming code provided in the Appendix. An informal argument as to the correctness of the approach comes from noting that the first derivative of say ll (β*) gives the estimating function where u (Y, Z, X; β*) = [1, Xi′, Zi]′ (Yi − Pr(Y = 1|Xi, Zi; β*)) is the logistic regression score function evaluated at β*, and ˆE (·|Yi, Zi, Si) stands for the conditional expectation with respect to ĝ, the estimated density function of Z given Y, X, and S; this estimating function is exactly an empirical version of the estimating function previously derived and used for efficiency purposes2,10 and subsequently shown to be doubly robust.6,7,8 The first term of the equation corresponds to inverse-probability weighting for logistic regression, and only uses complete cases. The role of the second term is to make optimal use of all of the observed data. It serves to recover information by leveraging any existing correlation between Zi and (Yi, Xi, Si) to “impute ” expected contributions to the estimating function for both complete and incomplete observations. When both models for π and g are correct, one can show that ˆβdr is generally asymptotically more efficient than the inverse-probability weighting estimator2 used by Moore et al,1 which solves σi=1nˆωiu(Yi, Zi, Xi; β*) = 0. A similar result holds for llK (β*), although ˆE (·|Yi, Zi, Si) is then replaced by a summation over the pseudo-observations Zi,k*, k = 1, . . . K. When Z is continuous, this summation is in effect evaluating an integral by Monte Carlo simulation, thus avoiding the need for numerical integration. The choice of K can only affect the efficiency of the resulting estimator and not its consistency. When all models are correct, smaller values of K will invariably lead to less efficiency than large ones; however, values of K between 5 and 15 should probably suffice in most applications. The estimator ˆβdr is asymptotically normal, therefore formal inferences on β can be based on Wald-type confidence intervals constructed with the nonparametric bootstrap estimator of variance. It is important to note that ll (β*) (equivalently llK (β*)) cannot be viewed as a proper log- likelihood function, but rather serves as a device to obtain an estimator with properties (a), (b), and (c). Finally, the proposed method can be applied to vector-valued Z with components consisting of continuous, discrete or covariates of mixed type, and is thus very general. A Simulation Study In a small simulation study aimed at illustrating the finite sample behavior of this estimator compared with inverse-probability weighting, I simulated 1000 random samples of size n = 500, of independent and identically distributed data under a generating mechanism which does the following: samples S and X from independent standard normal densities; generates dichotomous Z and Y from the joint probability mass function where I impose and so that logit (f(1|Z, X; β)) = β0 + βzZ + βxzZX + βxX and logit (g(1|Y, X, S; η)) = η0+ ηyY + ηyxYX + ηxX + ηsS, where ηY = βz and ηyx = βxz; and finally obtains R from Bernoulli sampling with success probability given by logit (π(X, Y, S; α)) = α0 + αyY + αsS + αxX + αyx2YX2 + αsySY + αs2xS2X, thus imposing MAR. The Table 1 reports results from fitting ˆβdr and ˆβipw (the inverse-probability–weighted estimator) to the observed data (RZ, R, X, Y, S) under 4 different model specifications; first, in estimating ˆβdr and ˆβipw I consistently estimate both π and g in the scenario corresponding to the “all true” row of results in the Table, next I mis-specify π by setting αyx2=αsy=αs2x = 0 under the setting which is reported in the “wrong π” row of results of the Table, then I mis-specify g by setting ηyz = 0 under the setting which is reported in the “wrong g” row of results of the Table, and finally both g and π are incorrectly specified as above, in the final row labeled “all wrong” results.TABLE 1: Simulation ResultsUnder the correct model specification of both π and g (corresponding to the first 2 rows of results in the Table), both estimators have finite sample bias of comparable size across parameter estimates. As predicted by theory, ˆβdr is more efficient than ˆβipw, with ranging between 87% and 92% for different parameters. Under mis-specification of g, both estimators maintain small finite sample bias, while neither appears to clearly dominate with respect to efficiency across all parameters. Next, I find that in agreement with theory, ˆβipw is severely biased under mis-specification of π relative to ˆβdr which again maintains negligible bias. Under mis-specification of both π and g, the bias of all estimators is also larger as expected. Survey Sampling I have briefly demonstrated a simple implementation of a doubly robust estimator of the parameters of a logistic regression model with covariates missing at random. Although this method is described under random sampling, it is easily adapted to the survey sampling setting. In fact, this is achieved by simply weighting each person's contribution to the pseudo-log-likelihood by the inverse of their probability of being selected into the sample. I revisited the NHANES III data used by Moore et al1 and obtained a doubly robust estimate of β using my approach. Remarkably, out of 14 logistic regression coefficients, my estimated effects agreed for all but 3 parameters; education: 12 = 0.282 (instead of 0.134 reported by Moore et al), >12 = 0.572 instead of (0.356); routine health care visit <2 years = 0.493 (instead of 0.703 reported by Moore et al); and history of premature heart disease = 0.202 (instead of 0.002). However, I did not obtain confidence intervals for these estimates, as the nonparametric bootstrap in this setting requires taking into account the design of the study, a task beyond the scope of this commentary. With very few changes, the approach proposed herein equally applies to linear and log-linear regression with missing covariates, most commonly used for continuous and count data respectively; details are omitted but are easily deduced from the accompanying technical report.9
- Research Article
86
- 10.1111/j.1468-0262.2004.00485.x
- Dec 10, 2003
- Econometrica
In this paper we propose a new estimator for a model with one endogenous regressor and many instrumental variables. Our motivation comes from the recent literature on the poor properties of standard instrumental variables estimators when the instrumental variables are weakly correlated with the endogenous regressor. Our proposed estimator puts a random coefficients structure on the relation between the endogenous regressor and the instruments. The variance of the random coefficients is modelled as an unknown parameter. In addition to proposing a new estimator, our analysis yields new insights into the properties of the standard two-stage least squares (TSLS) and limited-information maximum likelihood (LIML) estimators in the case with many weak instruments. We show that in some interesting cases, TSLS and LIML can be approximated by maximizing the random effects likelihood subject to particular constraints. We show that statistics based on comparisons of the unconstrained estimates of these parameters to the implicit TSLS and LIML restrictions can be used to identify settings when standard large sample approximations to the distributions of TSLS and LIML are likely to perform poorly. We also show that with many weak instruments, LIML confidence intervals are likely to have under-coverage, even though its finite sample distribution is approximately centered at the true value of the parameter. In an application with real data and simulations around this data set, the proposed estimator performs markedly better than TSLS and LIML, both in terms of coverage rate and in terms of risk.
- Single Report
42
- 10.1920/wp.cem.2011.2511
- Jun 25, 2011
This paper provides sufficient conditions for the nonparametric identification of the regression function m(.) in a regression model with an endogenous regressor x and an instrumental variable z. It has been shown that the identification of the regression function from the conditional expectation of the dependent variable on the instrument relies on the completeness of the distribution of the endogenous regressor conditional on the instrument, i.e., f(x/z). We provide sufficient conditions for the completeness of f(x/z) without imposing a specific functional form, such as the exponential family. We show that if the conditional density f(x/z) coincides with an existing complete density at a limit point in the support of z, then f(x/z) itself is complete, and therefore, the regression function m(.) is nonparametrically identified. We use this general result provide specific sufficient conditions for completeness in three different specifications of the relationship between the endogenous regressor x and the instrumental variable z.
- Research Article
14
- 10.1177/00222437231195577
- Nov 23, 2023
- Jmr, Journal of Marketing Research
Endogenous regressors can lead to biased estimates for causal effects using methods assuming regressor–error independence. To correct for endogeneity bias, the authors propose a new method that accounts for the regressor–error dependence using flexible semiparametric odds ratio conditional models; the approach requires neither parametric distributional assumptions nor tuning parameters for modeling endogenous regressors’ distributions conditional on the error term and exogenous regressors. Inference is achieved via optimizing the profile likelihood concentrating on the parameters of interest. The proposed approach requires no use of instrumental variables (IVs), observed or latent, that must satisfy the stringent condition of exclusion restriction. Nonnormally distributed endogenous regressors are required for model identification with a normal error distribution. The approach's flexibility in capturing regressor–error dependence increases the capability of IV-free endogeneity correction and provides opportunities to improve the accuracy of causal effect estimation. Unlike existing IV-free methods, the proposed approach can handle discrete endogenous regressors with few levels, such as binary regressors or count regressors with small means, and is thus applicable to a plethora of applications involving such regressors. The authors demonstrate the versatility of the approach for binary, count, and continuous endogenous regressors using comprehensive simulation studies and empirical data.
- Single Report
27
- 10.3386/t0204
- Sep 1, 1996
- National Bureau of Economic Research
In this paper, we explore Bayesian inference in models with many instrumental variables that are potentially weakly correlated with the endogenous regressor. The prior distribution has a hierarchical (nested) structure. We apply the methods to the Angrist-Krueger (AK, 1991) analysis of returns to schooling using instrumental variables formed by interacting quarter of birth with state/year dummy variables. Bound, Jaeger, and Baker (1995) show that randomly generated instrumental variables, designed to match the AK data set, give two-stage least squares results that look similar to the results based on the actual instrumental variables. Using a hierarchical model with the AK data, we find a posterior distribution for the parameter of interest that is tight and plausible. Using data with randomly generated instruments, the posterior distribution is diffuse. Most of the information in the AK data can in fact be extracted with quarter of birth as the single instrumental variable. Using artificial data patterned on the AK data, we find that if all the information had been in the interactions between quarter of birth and state/year dummies, then the hierarchical model would still have led to precise inferences, whereas the single instrument model would have suggested that there was no information in the data. We conclude that hierarchical modeling is a conceptually straightforward way of efficiently combining many weak instrumental variables.
- Research Article
1
- 10.3389/fsufs.2025.1657505
- Sep 3, 2025
- Frontiers in Sustainable Food Systems
The prevention and control of agricultural non-point source pollution (ANPP) is vital for promoting green agricultural development and has become a key public concern. Using statistical data from 30 provinces in China during 2005–2022, this study empirically investigates the influence mechanism and threshold effects of rural labor transfer on ANPP, employing a comprehensive framework including benchmark regression, mediating-moderating effect models, and threshold regression, supplemented by instrumental variable (IV) techniques to address endogeneity. Key findings include: (i) Rural labor transfer significantly exacerbates ANPP, with heterogeneous effects across rural labor transfer types and regional contexts. (ii) A single threshold effect exists, demonstrating a non-linear pattern where the marginal impact of rural labor transfer on ANPP diminishes as its scale increases. (iii) Non-farm income serves as a critical mediating pathway through which rural labor transfer intensifies ANPP. (iv) Agricultural socialized services moderate this effect by mitigating the mediating role of non-farm income, thereby alleviating environmental degradation. These findings provide policy insights for China to optimize ANPP prevention strategies, highlighting the need to coordinate labor migration management, income diversification, and agricultural socialized services development for sustainable agricultural growth.
- Research Article
30
- 10.1016/j.jeconom.2019.07.006
- Sep 5, 2019
- Journal of Econometrics
Panel threshold regressions with latent group structures
- Research Article
5
- 10.1016/j.econlet.2018.08.039
- Sep 5, 2018
- Economics Letters
Threshold regression asymptotics: From the compound Poisson process to two-sided Brownian motion
- Report Series
409
- 10.1920/wp.cem.2001.0901
- Nov 1, 2001
- Cemmap working papers
This paper considers the nonparametric and semiparametric methods for estimating regression models with continuous endogenous regressors. We list a number of different generalizations of the linear structural equation model, and discuss how two common estimation approaches for linear equations-the "instrumental variables" and "control function" approaches-may be extended to nonparametric generalizations of the linear model and to their semiparametric variants. We consider the identification and estimation of the "Average Structural Function" and argue that this is a parameter of central interest in the analysis of semiparametric and nonparametric models with endogenous regressors. We consider a particular semiparametric model, the binary response model with linear index function and nonparametric error distribution, and describes in detail how estimation of the parameters of interest can be constructed using the "control function" approach. This estimator is applied to estimating the relation of labor force participation to nonlabor income, viewed as an endogenous regressor.
- Supplementary Content
- 10.6092/unibo/amsdottorato/8813
- Apr 9, 2019
- AMS Dottorato Institutional Doctoral Theses Repository (University of Bologna)
Instrumental Variables (IV) are widely used in econometrics to overcome endogeneity problem in regression models, which occurs when regressors are correlated with the stochastic component. Nonetheless, in applied works, practitioners face with instruments that are collectively ``weak'', i.e. poorly correlated with endogenous regressors. Under weak instruments, conventional estimators are no longer consistent and asymptotically normal. Furthermore, bootstrap methods could be useful to improve inference in IV estimation. However, under poorly relevant instruments, the bootstrap is deemed invalid and its use is generally discouraged in applied papers. In this work, we propose a new derivation of bootstrapped IV estimators under weak instruments asymptotics (Stock and Yogo, 2005) using residual--based bootstrap method involving fixed or resampled instruments. We prove that bootstrap counterpart of estimators, conditionally on the data, converges to a random distribution preserving some patterns (non--normality) of weak and irrelevant instruments scenarios. These issues may be also reflected in bootstrap--based confidence sets and hypothesis testing. In this sense, we explore the usefulness of bootstrap methods to provide information on the weakness (or the strength) of the instruments. We consider descriptive indicators and develop new bootstrap-based tests useful to detect weak instruments in IV framework. The method basically relies on Angelini et al. (2016) and allows to test normality of a certain number of (possibly standardized) bootstrap replications. Since conventional normality tests can lose power in presence of more instruments and high endogeneity, we propose new test statistics with the aim to test standard normality on the bootstrap replications. These tests are based on the moments of standard normal and are asymptotically chi--square distributed under the null hypothesis. In conclusion, we find that, in some cases, bootstrapped estimators may be used to test weak identification.
- Research Article
177
- 10.1017/s0266466609990727
- Mar 17, 2010
- Econometric Theory
We consider estimation of parameters in a regression model with endogenous regressors. The endogenous regressors along with a large number of other endogenous variables are driven by a small number of unobservable exogenous common factors. We show that the estimated common factors can be used as instrumental variables and they are more efficient than the observed variables in our framework. Whereas standard optimal generalized method of moments estimator using a large number of instruments is biased and can be inconsistent, the factor instrumental variable estimator (FIV) is shown to be consistent and asymptotically normal, even if the number of instruments exceeds the sample size. Furthermore, FIV remains consistent even if the observed variables are invalid instruments as long as the unobserved common components are valid instruments. We also consider estimating panel data models in which all regressors are endogenous but share exogenous common factors. We show that valid instruments can be constructed from the endogenous regressors. Although single equation FIV requires no bias correction, the faster convergence rate of the panel estimator is such that a bias correction is necessary to obtain a zero-centered normal distribution.
- Research Article
312
- 10.1162/rest_a_00171
- Jun 27, 2008
- Review of Economics and Statistics
Dealing with endogenous regressors is a central challenge of applied research. The standard solution is to use instrumental variables that are assumed to be uncorrelated with unobservables. We instead assume (i) the correlation between the instrument and the error term has the same sign as the correlation between the endogenous regressor and the error term, and (ii) that the instrument is less correlated with the error term than is the endogenous regressor. Using these assumptions, we derive analytic bounds for the parameters. We demonstrate the method in two applications.
- Research Article
- 10.1515/snde-2012-0071
- Jan 1, 2014
- Studies in Nonlinear Dynamics & Econometrics
This paper extends Kim’s (Kim, C.-J. 2004. “Markov-Switching Models with Endogenous Explanatory Variables.”