Econometrics, economics, finance, random rants.

Econometrics, economics, finance, random rants...
Showing posts with label Causal modeling. Show all posts
Showing posts with label Causal modeling. Show all posts

Wednesday, April 10, 2019

Bad News for IV Estimation

Alwyn Young has an eye-opening recent paper, "Consistency without Inference: Instrumental Variables in Practical Application".  There's a lot going on worth thinking about in his Monte Carlo:  OLS vs. IV; robust/clustered s.e.'s vs. not; testing/accounting for weak instruments vs. not; jacknife/bootstrap vs. "conventional" inference; etc.  IV as typically implemented comes up looking, well, dubious.

Alwyn's related analysis of published studies is even more striking.  He shows that, in a sample of 1359 IV regressions in 31 papers published in the journals of the American Economic Association,
"... statistically significant IV results generally depend upon only one or two observations or clusters, excluded instruments often appear to be irrelevant, there is little statistical evidence that OLS is actually substantively biased, and IV confidence intervals almost always include OLS point estimates." 
Wow.

Perhaps the high leverage is Alwyn's most striking result, particularly as many empirical economists seem to have skipped class on the day when leverage assessment was taught.  Decades ago, Marjorie Flavin attempted some remedial education in her 1991 paper, "The Joint Consumption/Asset Demand Decision: A Case Study in Robust Estimation".  She concluded that
"Compared to the conventional results, the robust instrumental variables estimates are more stable across different subsamples, more consistent with the theoretical specification of the model, and indicate that some of the most striking findings in the conventional results were attributable to a single, highly unusual observation." 
Sound familiar?  The non-robustness of conventional IV seems disturbingly robust, from Flavin to Young.

Flavin's paper evidently fell on deaf ears and remains unpublished. Hopefully Young's will not meet the same fate.

Monday, April 8, 2019

Identification via the ZLB and More

Sophocles Mavroeidis at Oxford has a very nice paper on using the nominal interest rate zero lower bound (ZLB) to identify VAR's.  Effectively, hitting the ZLB is a form of (endogenous) structural change that can be exploited for identification.  He has results showing whether/when one has point identification, set identification, or no identification. Really good stuff.

An interesting question is whether there may be SETS of bounds that may be hit. Suppose so, and suppose that we don't know whether/when they'll be hit, but we do know that if/when one bound is hit, all bounds are hit. An example might be nominal short rates in two countries with tightly-integrated money markets.

Now recall the literature on testing for multivariate structural change, which reveals large power increases in such situations (Bai, Lumsdaine and Stock). In Sophocles' case, it suggests the potential for greatly sharpened set ID.  Of course it all depends on the truth/relevance of my supposition...




Monday, March 25, 2019

Ensemble Methods for Causal Prediction

Great to see ensemble learning methods (i.e., forecast combination) moving into areas of econometrics beyond time series / macro-econometrics, where they have thrived ever since Bates and Granger (1969), generating a massive and vibrant literature.  (For a recent contribution, including historical references, see Diebold and Shin, 2019.)  In particular, the micro-econometric / panel / causal literature is coming on board.  See for example this new and interesting paper by Susan Athey et al.

Sunday, December 16, 2018

Causality as Robust Prediction

I like thinking about causal estimation as a type of prediction (e.g., here). Here's a very nice slide deck from Peter Buhlmann at ETH Zurich detailing his group's recent and ongoing work in that tradition.














Saturday, September 29, 2018

RCT's vs. RDD's

Art Owen and Hal Varian have an eye-opening new paper, "Optimizing the Tie-Breaker Regression Discontinuity Design".

Randomized controlled trials (RCT's) are clearly the gold standard in terms of statistical efficiency for teasing out causal effects. Assume that you really can do an RCT. Why then would you ever want to do anything else?

Answer: There may be important considerations beyond statistical efficiency. Take the famous "scholarship example". (You want to know whether receipt of an academic scholarship causes enhanced academic performance among strong scholarship test performers.) In an RCT approach you're going to give lots of academic scholarships to lots of randomly-selected people, many of whom are not strong performers. That's wasteful. In a regression discontinuity design (RDD) approach ("give scholarships only to strong performers who score above X in the scholarship exam, and compare the performances of students who scored just above and below X"), you don't give any scholarships to weak performers. So it's not wasteful -- but the resulting inference is statistically inefficient. 

"Tie breakers" implement a middle ground: Definitely don't give scholarships to bottom performers, definitely do give scholarships to top performers, and randomize for a middle group. So you gain some efficiency relative to pure RDD (but you're a little wasteful), and you're less wasteful than a pure RCT (but you lose some efficiency).

Hence there's an trade-off, and your location on it depends on the size of the your middle group. Owen and Varian characterize the trade-off and show how to optimize the size of the middle group. Really nice, clean, and useful.

[Sorry but I'm running way behind. I saw Hal present this work a few months ago at a fine ECB meeting on predictive modeling.]

Sunday, July 30, 2017

Regression Discontinuity and Event Studies in Time Series

Check out the new paper, "Regression Discontinuity in Time [RDiT]: Considerations for Empirical Applications", by Catherine Hausman and David S. Rapson.  (NBER Working Paper No. 23602, July 2017.  Ungated copy here.)

It's interesting in part because it documents and contributes to the largely cross-section regression discontinuity design literature's awakening to time series. But the elephant in the room is the large time-series "event study" (ES) literature, mentioned but not emphasized by Hausman and Rapson.  [In a one-sentence nutshell, here's how an ES works: model the pre-event period, use the fitted pre-event model to predict the post-event period, and ascribe any systematic forecast error to the causal impact of the event.]  ES's trace to the classic Fama et al. (1969).  Among many others, MacKinlay's 1997 overview is still fresh, and Gürkaynak and Wright (2013) provide additional perspective.

One question is what the RDiT approach adds to the ES approach, and related, what it adds to well-developed time-series toolkit of other methods for assessing structural change. At present, and notwithstanding the Hausman-Rapson paper, my view is "little or nothing".  Indeed in most respects it would seem that a RDiT study *is* an ES, and conversely.  So call it what you will, "ES" or "RDiT"

But there are important open issues in ES / RDiT, and Hausman-Rapson correctly emphasize one of them, namely issues and difficulties associated with "wide" pre- and post-event windows, which is often the relevant case in time series.

Things are generally "easy" in cross sections, where we can usually take narrow windows (e.g., in the classic scholarship exam example, we use only test scores very close to the scholarship threshold).  Things are similarly "easy" in time series *IF* we can take similarly narrow windows (e.g., high-frequency asset return data facilitate taking narrow pre- and post-event windows in financial applications).  In such cases it's comparatively easy to credibly ascribe a post-event break to the causal impact of the event.

But in other time-series areas like macro and environmental, we might want (or need) to use wide pre- and post-event windows.  Then the trick becomes modeling the pre- and post-event periods successfully enough so that we can credibly assert that any structural change is due exclusively to the event -- very challenging, but not hopeless.

Hats off to Hausman and Rapson for beginning to bridge the ES and regression discontinuity literatures, and for implicitly helping to push the ES literature forward.

Tuesday, July 25, 2017

Time-Series Regression Discontinuity

I'll have something to say in next week's post.  Meanwhile check out the interesting new paper, "Regression Discontinuity in Time: Considerations for Empirical Applications", by Catherine Hausman and David S. Rapson, NBER Working Paper No. 23602, July 2017.  (Ungated version here.)

Sunday, February 19, 2017

Econometrics: Angrist and Pischke are at it Again

Check out the new Angrist-Pischke (AP), "Undergraduate Econometrics Instruction: Through Our Classes, Darkly".

I guess I have no choice but to weigh in. The issues are important, and my earlier AP post, "Mostly Harmless Econometrics?", is my all-time most popular.

Basically AP want all econometrics texts to look a lot more like theirs. But their books and their new essay unfortunately miss (read: dismiss) half of econometrics.

Here's what AP get right:

(Goal G1) One of the major goals in econometrics is predicting the effects of exogenous "treatments" or "interventions" or "policies". Phrased in the language of estimation, the question is "If I intervene and give someone a certain treatment \({\partial x}, x \in X\), what is my minimum-MSE estimate of her \(\ \partial y\)?" So we are estimating the partial derivative \({\partial y / \partial x}\).

AP argue the virtues and trumpet the successes of a "design-based" approach to G1. In my view they make many good points as regards G1: discontinuity designs, dif-in-dif designs, and other clever modern approaches for approximating random experiments indeed take us far beyond "Stones'-age" approaches to G1. 
(AP sure turn a great phrase...). And the econometric simplicity of the design-based approach is intoxicating: it's mostly just linear regression of \(y\) on \(x\) and a few cleverly-chosen control variables -- you don't need a full model -- with White-washed standard errors. Nice work if you can get it. And yes, moving forward, any good text should feature a solid chapter on those methods.

Here's what AP miss/dismiss:

(Goal G2) The other major goal in econometrics is predicting \(y\). In the language of estimation, the question is "If a new person \(i\) arrives with covariates \(X_i\), what is my minimum-MSE estimate of her \(y_i\)? So we are estimating a conditional mean \(E(y | X) \), which in general is very different from estimating a partial derivative \({\partial y / \partial x}\).

The problem with the AP paradigm is that it doesn't work for goal G2. Modeling nonlinear functional form is important, as the conditional mean function \(E(y | X) \) may be highly nonlinear in \(X\); systematic model selection is important, as it's not clear a priori what subset of \(X\) (i.e., what model) might be most useful for approximating \(E(y | X) \); detecting and modeling heteroskedasticity is important (in both cross sections and time series), as it's the key to accurate interval and density prediction; detecting and modeling serial correlation is crucially important in time-series contexts, as "the past" is the key conditioning information for predicting "the future"; etc., etc, ... 


(Notice how often "model" and "modeling" appear in the above paragraph. That's precisely what AP dismiss, even in their abstract, which very precisely, and incorrectly, declares that "Applied econometrics ...[now prioritizes]... the estimation of specific causal effects and empirical policy analysis over general models of outcome determination".)

The AP approach to goal G2 is to ignore it, in a thinly-veiled attempt to equate econometrics exclusively with G1, which nicely feathers the AP nest. Sorry guys, but no one's buying it. That's why the textbooks continue to feature G2 tools and techniques so prominently, as well they should.



Tuesday, January 3, 2017

Torpedoing Econometric Randomized Controlled Trials

A very Happy New Year to all!

I get no pleasure from torpedoing anything, and "torpedoing" is likely exaggerated, but nevertheless take a look at "A Torpedo Aimed Straight at HMS Randomista". It argues that many econometric randomized controlled trials (RCT's) are seriously flawed -- not even internally valid -- due to their failure to use double-blind randomization. At first the non-double-blind critique may sound cheap and obvious, inviting you to roll your eyes and say "get over it". But ultimately it's not.

Note the interesting situation. Everyone these days is worried about external validity (extensibility), under the implicit assumption that internal validity has been achieved (e.g., see this earlier post). But the 
non-double-blind critique makes clear that even internal validity may be dubious in econometric RCT's as typically implemented.

The underlying research paper, "Behavioural Responses and the Impact of New Agricultural Technologies: Evidence from a Double-Blind Field Experiment in Tanzania", by Bulte et al., was published in 2014 in the American Journal of Agricultural Economics. Quite an eye-opener
.

Here's the abstract:

Randomized controlled trials in the social sciences are typically not double-blind, so participants know they are “treated” and will adjust their behavior accordingly. Such effort responses complicate the assessment of impact. To gauge the potential magnitude of effort responses we implement an open RCT and double-blind trial in rural Tanzania, and randomly allocate modern and traditional cowpea seed-varieties to a sample of farmers. Effort responses can be quantitatively important––for our case they explain the entire “treatment effect on the treated” as measured in a conventional economic RCT. Specifically, harvests are the same for people who know they received the modern seeds and for people who did not know what type of seeds they got, but people who knew they received the traditional seeds did much worse. We also find that most of the behavioral response is unobserved by the analyst, or at least not readily captured using coarse, standard controls.

Sunday, December 11, 2016

Varieties of RCT Extensibility

Even internally-valid RCT's have issues. They reveal the treatment effect only for the precise experiment performed and situation studied. Consider, for example, a study of the effects of fertilizer on crop yield, done for region X during a heat wave. Even if internally valid, the estimated treatment effect is that of fertilizer on crop yield in region X during a heat wave. The results do not necessarily generalize -- and in this example surely do not generalize -- to times of ``normal" weather, even in region X. And of course, for a variety of reasons, they may not generalize to regions other than X, even in heat waves.

Note the interesting time-series dimension to the failure of external validity (extensibility) in the example above. (The estimate is obtained during this year's heat wave, but next year may be "normal", or "cool". And this despite the lack of any true structural change. But of course there could be true structural change, which would only make matters worse.) This contrasts with the usual cross-sectional focus of extensibility discussions (e.g., we get effect e in region X, but what effect would we get in region Z?)

In essence, we'd like panel data, to account both for cross-section effects and time-series effects, but most RCT's unfortunately have only a single cross section.

Mark Rosenzweig and Chris Udry have a fascinating new paper, "Extenal Validity in a Stochastic World", that grapples with some of the time-series extensibility issues raised above.

Sunday, July 3, 2016

DAG Software

Some time ago I mentioned the DAG (directed acyclical graph) primer by Judea Pearl et al.  As noted in Pearl's recent blog post, a manual will be available with software solutions based on a DAGitty R package.  See http://dagitty.net/primer/

More generally -- that is, quite apart from the Pearl et al. primer -- check out DAGity at http://dagitty.net.  Click on "launch" and play around for a few minutes. Very cool. 

Tuesday, June 21, 2016

Conditional Dependence and Partial Correlation

In the multivariate normal case, conditional independence is the same as zero partial correlation.  (See below.) That makes a lot of things a lot simpler.  In particular, determining ordering in a DAG is just a matter of assessing partial correlations. Of course in many applications normality may not hold, but still...

Aust. N.Z. J. Stat. 46(4), 2004, 657–664
PARTIAL CORRELATION AND CONDITIONAL CORRELATION AS MEASURES OF CONDITIONAL INDEPENDENCE
Kunihiro Baba1∗, Ritei Shibata1 and Masaaki Sibuya2
Keio University and Takachiho University
Summary
This paper investigates the roles of partial correlation and conditional correlation as mea-sures of the conditional independence of two random variables. It first establishes a suffi-cientconditionforthecoincidenceofthepartialcorrelationwiththeconditionalcorrelation. The condition is satisfied not only for multivariate normal but also for elliptical, multi-variate hypergeometric, multivariate negative hypergeometric, multinomial and Dirichlet distributions. Such families of distributions are characterized by a semigroup property as a parametric family of distributions. A necessary and sufficient condition for the coinci-dence of the partial covariance with the conditional covariance is also derived. However, a known family of multivariate distributions which satisfies this condition cannot be found, except for the multivariate normal. The paper also shows that conditional independence has no close ties with zero partial correlation except in the case of the multivariate normal distribution; it has rather close ties to the zero conditional correlation. It shows that the equivalence between zero conditional covariance and conditional independence for normal variables is retained by any monotone transformation of each variable. The results suggest that care must be taken when using such correlations as measures of conditional indepen-dence unless the joint distribution is known to be normal. Otherwise a new concept of conditional independence may need to be introduced in place of conditional independence through zero conditional correlation or other statistics.
Keywords: elliptical distribution; exchangeability; graphical modelling; monotone transformation.

Friday, June 3, 2016

Causal Estimation and Millions of Lives

This just in from a fine former Ph.D. student.  He returned to India many years ago and made his fortune in finance.  He's now devoting himself the greater good, working with the Bill and Melinda Gates Foundation.

I reminded him that I'm not likely to be a big help, as I generally don't do causal estimation or experimental design. But he kindly allowed me to post his communication below (abridged and slightly edited). Please post comments for him if you have any suggestions. [As you know, I write this blog more like a newspaper column, neither encouraging nor receiving many comments -- so now's your chance to comment!]

He writes:

One of the key challenges we face in our work is that causality is not known, and while theory and large scale studies, such as those published in the Lancet, do provide us with some guidance, it is far from clear that they reflect the reality on the ground when we are intervening in field settings with markedly different starting points from those that were used in the studies. However, while we observe the ground situation imperfectly and with large error, the inertia in the underlying system that we are trying to impact is so high that that it would perhaps be safe to say that, unlike in the corporate world, there isn’t a lot of creative destruction going on here. In such a situation it would seem to me that the best way to learn about the “true but unobserved” reality and how to permanently change it and scale the change cost-effectively (such as nurse behavior in facilities) is to go on attempting different interventions which are structured in such a way as to allow for a rapid convergence to the most effective interventions (similar to the famous Runge-Kutta iterative methods for rapidly and efficiently arriving at solutions to differential equations to the desired level of accuracy).

However, while the need is for rapid learning, the most popular methods proceed by collecting months or years of data in both intervention and control settings, and at the end of it all, if done very-very carefully, all that they can tell you is that there were some links (or not) between the interventions and results without giving you any insight into why something happened or what can be done to improve it. In the meanwhile one is expected to hold the intervention steady and almost discard all the knowledge that is continuously being generated and be patient even while lives are being lost because the intervention was not quite designed well. While the problems with such an approach are apparent, the alternative cannot be instinct or gut feeling and a series of uncoordinated actions in the name of “being responsive”.

I am writing to request your help in pointing us to literature that can act as a guide to how we may do this better. ... I have indeed found some ideas in the literature that may be somewhat useful, ... [and] while very interesting and informative, I’m afraid it is not yet clear to me how we will apply these ideas in our actual field settings, and how we will design our Measurement, Learning, and Evaluation approaches differently so that we can actually implement these ideas in difficult on-ground settings in remote parts of our country involving, literally, millions of lives.

Sunday, April 10, 2016

On "The Human Capital Approach to Inference"

Check out the interesting new paper by Bentley MacLeod at Columbia ("The Human Capital Approach to Inference"), on using economic theory in combination with machine learning to estimate conditional average treatment effects better than can be done with randomized control trials.

Quite apart from new methods for accurate estimation of conditional average treatment effects, the paper's intro contains some interesting tidbits on causal econometric inference. Here's one sequence in yellow, with my reactions:

BM: "There are two distinct approaches to modern empirical economics."
-- The MacLeod paper is exclusively about causal inference, so it should say "two distinct approaches to causal inference in modern empirical economics." Equating causal inference to all of empirical economics is simply wrong. Causal inference is a large and very important part of modern empirical economics, but far from its entirety. The booming field of financial econometrics, for example, is largely and intentionally reduced-form. See this.

BM: "First, there is research using structural models that begins by assuming individuals make utility maximizing decisions within a well defined environment, and then proceeds to measure the value of the unknown parameters..."
-- There is some unsettling truth here. A cynical but not-entirely-false view is that structural causal inference effectively assumes a causal mechanism, known up to a vector of parameters that can be estimated. Big assumption. And of course different structural modelers can make different assumptions and get different results.

BM: "The second approach addresses the self-selection of individuals into different observed treatments or choices by either explicitly randomizing treatments/choices in the context of an experiment...or through the use of a natural experiment that allows for an instrumental variables strategy. There is general agreement that explicit randomization provides one of the cleanest ways to obtain a measure of the effect of choice."
-- There's rarely general agreement about anything in economics. But yes, randomization is arguably the gold standard for causal effect estimation, if and when it can be done credibly.

Monday, January 19, 2015

Haavelmo and Causal Modeling

Speaking of causal modeling, as we were in a recent postEconometric Theory is doing a special issue on Haavelmo. (This is not new information, but hey, I'm usually slow to notice things, and perhaps you are too.) Should be a fine issue, with many fascinating papers, including Heckman-Pinto. I list some of them below, from ET's "First View" page.

HAAVELMO’S CONTRIBUTIONS TO SIMULTANEOUS-EQUATIONS ESTIMATION