[Click on "Machine Learning" at right for earlier "Machine Learning and Econometrics" posts.]
We econometricians need -- and have always had -- cross section and time series ("micro econometrics" and "macro/financial econometrics"), causal estimation and predictive modeling, structural and non-structural. And all continue to thrive.
But there's a new twist, happening now, making this an unusually exciting time in econometrics. Predictive econometric modeling is not only alive and well, but also blossoming anew, this time at the interface of micro-econometrics and machine learning. A fine example is the new Kleinberg, Lakkaraju, Leskovic, Ludwig and Mullainathan paper, “Human Decisions and Machine Predictions”, NBER Working Paper 23180 (February 2017).
Good predictions promote good decisions, and econometrics is ultimately about helping people to make good decisions. Hence the new developments, driven by advances in machine learning, are most welcome contributions to a long and distinguished predictive econometric modeling tradition.
Econometrics, economics, finance, random rants.
Econometrics, economics, finance, random rants...
Showing posts with label Cross-section econometrics. Show all posts
Showing posts with label Cross-section econometrics. Show all posts
Sunday, March 19, 2017
Monday, March 13, 2017
ML and Metrics VII: Cross-Section Non-Linearities
[Click on "Machine Learning" at right for earlier "Machine Learning and Econometrics" posts.]
The predictive modeling perspective needs not only to be respected and embraced in econometrics (as it routinely is, notwithstanding the Angrist-Pischke revisionist agenda), but also to be enhanced by incorporating elements of statistical machine learning (ML). This is particularly true for cross-section econometrics insofar as time-series econometrics is already well ahead in that regard. For example, although flexible non-parametric ML approaches to estimating conditional-mean functions don't add much to time-series econometrics, they may add lots to cross-section econometric regression and classification analyses, where conditional mean functions may be highly nonlinear for a variety of reasons. Of course econometricians are well aware of traditional non-parametric issues/approaches, especially kernel and series methods, and they have made many contributions, but there's still much more to be learned from ML.
The predictive modeling perspective needs not only to be respected and embraced in econometrics (as it routinely is, notwithstanding the Angrist-Pischke revisionist agenda), but also to be enhanced by incorporating elements of statistical machine learning (ML). This is particularly true for cross-section econometrics insofar as time-series econometrics is already well ahead in that regard. For example, although flexible non-parametric ML approaches to estimating conditional-mean functions don't add much to time-series econometrics, they may add lots to cross-section econometric regression and classification analyses, where conditional mean functions may be highly nonlinear for a variety of reasons. Of course econometricians are well aware of traditional non-parametric issues/approaches, especially kernel and series methods, and they have made many contributions, but there's still much more to be learned from ML.
Monday, June 6, 2016
Fixed Effects Without Panel Data
Consider a pure cross section (CS) of size N. Generally you'd like to allow for individual effects, but you can't, because OLS with a full set of N individual dummies is conceptually infeasible. (You'd exhaust degrees of freedom.) That's usually what motivates the desirability/beauty of panel data -- there you have NxT observations, so including N individual dummies becomes conceptually feasible.
But there's no need to stay with OLS. You can recover d.f. using regularization estimators like ridge (shrinkage) or LASSO (shrinkage and selection). So including a full set of individual dummies, even in a pure CS, is completely feasible! For implementation you just have to select the ridge or lasso penalty parameter, which is reliably done by cross validation (say).
There are two key points. The first is that you can allow for individual fixed effects even in a pure CS; that is, there's no need for panel data. That's what I've emphasized so far.
The second is that the proposed method actually gives estimates of the fixed effects. Sometimes they're just nuisance parameters that can be ignored; indeed standard panel estimation methods "difference them out", so they're not even estimated. But estimates of the fixed effects are crucial for forecasting: to forecast y_i, you need not only Mr. i's covariates and estimates of the "slope parameters", but also an estimate of Mr. i's intercept! That's why forecasting is so conspicuously absent from most of the panel literature -- the fixed effects are not estimated, so forecasting is hopeless. Regularized estimation, in contrast, delivers estimates of fixed effects, thereby facilitating forecasting, and you don't even need a panel.
But there's no need to stay with OLS. You can recover d.f. using regularization estimators like ridge (shrinkage) or LASSO (shrinkage and selection). So including a full set of individual dummies, even in a pure CS, is completely feasible! For implementation you just have to select the ridge or lasso penalty parameter, which is reliably done by cross validation (say).
There are two key points. The first is that you can allow for individual fixed effects even in a pure CS; that is, there's no need for panel data. That's what I've emphasized so far.
The second is that the proposed method actually gives estimates of the fixed effects. Sometimes they're just nuisance parameters that can be ignored; indeed standard panel estimation methods "difference them out", so they're not even estimated. But estimates of the fixed effects are crucial for forecasting: to forecast y_i, you need not only Mr. i's covariates and estimates of the "slope parameters", but also an estimate of Mr. i's intercept! That's why forecasting is so conspicuously absent from most of the panel literature -- the fixed effects are not estimated, so forecasting is hopeless. Regularized estimation, in contrast, delivers estimates of fixed effects, thereby facilitating forecasting, and you don't even need a panel.
Friday, June 3, 2016
Causal Estimation and Millions of Lives
This just in from a fine former Ph.D. student. He returned to India many years ago and made his fortune in finance. He's now devoting himself the greater good, working with the Bill and Melinda Gates Foundation.
I reminded him that I'm not likely to be a big help, as I generally don't do causal estimation or experimental design. But he kindly allowed me to post his communication below (abridged and slightly edited). Please post comments for him if you have any suggestions. [As you know, I write this blog more like a newspaper column, neither encouraging nor receiving many comments -- so now's your chance to comment!]
He writes:
I reminded him that I'm not likely to be a big help, as I generally don't do causal estimation or experimental design. But he kindly allowed me to post his communication below (abridged and slightly edited). Please post comments for him if you have any suggestions. [As you know, I write this blog more like a newspaper column, neither encouraging nor receiving many comments -- so now's your chance to comment!]
He writes:
One of the key challenges we face in our work is that causality is not known, and while theory and large scale studies, such as those published in the Lancet, do provide us with some guidance, it is far from clear that they reflect the reality on the ground when we are intervening in field settings with markedly different starting points from those that were used in the studies. However, while we observe the ground situation imperfectly and with large error, the inertia in the underlying system that we are trying to impact is so high that that it would perhaps be safe to say that, unlike in the corporate world, there isn’t a lot of creative destruction going on here. In such a situation it would seem to me that the best way to learn about the “true but unobserved” reality and how to permanently change it and scale the change cost-effectively (such as nurse behavior in facilities) is to go on attempting different interventions which are structured in such a way as to allow for a rapid convergence to the most effective interventions (similar to the famous Runge-Kutta iterative methods for rapidly and efficiently arriving at solutions to differential equations to the desired level of accuracy).
However, while the need is for rapid learning, the most popular methods proceed by collecting months or years of data in both intervention and control settings, and at the end of it all, if done very-very carefully, all that they can tell you is that there were some links (or not) between the interventions and results without giving you any insight into why something happened or what can be done to improve it. In the meanwhile one is expected to hold the intervention steady and almost discard all the knowledge that is continuously being generated and be patient even while lives are being lost because the intervention was not quite designed well. While the problems with such an approach are apparent, the alternative cannot be instinct or gut feeling and a series of uncoordinated actions in the name of “being responsive”.
I am writing to request your help in pointing us to literature that can act as a guide to how we may do this better. ... I have indeed found some ideas in the literature that may be somewhat useful, ... [and] while very interesting and informative, I’m afraid it is not yet clear to me how we will apply these ideas in our actual field settings, and how we will design our Measurement, Learning, and Evaluation approaches differently so that we can actually implement these ideas in difficult on-ground settings in remote parts of our country involving, literally, millions of lives.
However, while the need is for rapid learning, the most popular methods proceed by collecting months or years of data in both intervention and control settings, and at the end of it all, if done very-very carefully, all that they can tell you is that there were some links (or not) between the interventions and results without giving you any insight into why something happened or what can be done to improve it. In the meanwhile one is expected to hold the intervention steady and almost discard all the knowledge that is continuously being generated and be patient even while lives are being lost because the intervention was not quite designed well. While the problems with such an approach are apparent, the alternative cannot be instinct or gut feeling and a series of uncoordinated actions in the name of “being responsive”.
I am writing to request your help in pointing us to literature that can act as a guide to how we may do this better. ... I have indeed found some ideas in the literature that may be somewhat useful, ... [and] while very interesting and informative, I’m afraid it is not yet clear to me how we will apply these ideas in our actual field settings, and how we will design our Measurement, Learning, and Evaluation approaches differently so that we can actually implement these ideas in difficult on-ground settings in remote parts of our country involving, literally, millions of lives.
Saturday, October 17, 2015
Athey and Imbens on Machine Learning and Econometrics
Check out Susan Athey and Guido Imbens' NBER Summer Institute 2015 "Lectures on Machine Learning". (Be sure to scroll down, as there are four separate videos.) I missed the lectures this summer, and I just remembered that they're on video. Great stuff, reflecting parts of an emerging blend of machine learning (ML), time-series econometrics (TSE) and cross-section econometrics (CSE).
The characteristics of ML are basically (1) emphasis on overall modeling, for prediction (as opposed, for example, to emphasis on inference), (2) moreover, emphasis on non-causal modeling and prediction, (3) emphasis on computationally-intensive methods and algorithmic development, and (4) emphasis on large and often high-dimensional datasets.
Readers of this blog will recognize the ML characteristics as closely matching those of TSE! Rob Engle's V-Lab at NYU Stern's Volatility Institute, for example, embeds all of (1)-(4). So TSE and ML have a lot to learn from each other, but the required bridge is arguably quite short.
Interestingly, Athey and Imbens come not from the TSE tradition, but rather from the CSE tradition, which typically emphasizes causal estimation and inference. That makes for a longer required CSE-ML bridge, but it may also make for a larger payoff from building and crossing it (in both directions).
In any event I share Athey and Imbens' excitement, and I welcome any and all cross-fertilization of ML, TSE and CSE.
The characteristics of ML are basically (1) emphasis on overall modeling, for prediction (as opposed, for example, to emphasis on inference), (2) moreover, emphasis on non-causal modeling and prediction, (3) emphasis on computationally-intensive methods and algorithmic development, and (4) emphasis on large and often high-dimensional datasets.
Readers of this blog will recognize the ML characteristics as closely matching those of TSE! Rob Engle's V-Lab at NYU Stern's Volatility Institute, for example, embeds all of (1)-(4). So TSE and ML have a lot to learn from each other, but the required bridge is arguably quite short.
Interestingly, Athey and Imbens come not from the TSE tradition, but rather from the CSE tradition, which typically emphasizes causal estimation and inference. That makes for a longer required CSE-ML bridge, but it may also make for a larger payoff from building and crossing it (in both directions).
In any event I share Athey and Imbens' excitement, and I welcome any and all cross-fertilization of ML, TSE and CSE.
Subscribe to:
Posts (Atom)