Econometrics, economics, finance, random rants.

Econometrics, economics, finance, random rants...
Showing posts with label Model Selection. Show all posts
Showing posts with label Model Selection. Show all posts

Monday, May 8, 2017

Replicating Anomalies

I blogged a few weeks ago on "the file drawer problem".  In that vein, check out the interesting new paper below. I like their term "p-hacking". 

Random thought 1:  
Note that reverse p-hacking can also occur, when an author wants low p-values.  In the study below, for example, the deck could be stacked with all sorts of dubious/spurious "anomaly variables" that no one ever took seriously.  Then of course a very large number would wind up with low p-values.  I am not suggesting that the study below is guilty of this; rather, I simply had never thought about reverse p-hacking before, and this paper led me to think of the possibility, so I'm relaying the thought.

Related random thought 2:  
It would be interesting to compare anomalies published in "top journals" and "non-top journals" to see whether the top journals are more guilty or less guilty of p-hacking.  I can think of competing factors that could tip it either way!

Replicating Anomalies
by Kewei Hou, Chen Xue, Lu Zhang - NBER Working Paper #23394
Abstract:
The anomalies literature is infested with widespread p-hacking. We replicate the entire anomalies literature in finance and accounting by compiling a largest-to-date data library that contains 447 anomaly variables. With microcaps alleviated via New York Stock Exchange breakpoints and value-weighted returns, 286 anomalies (64%) including 95 out of 102 liquidity variables (93%) are insignificant at the conventional 5% level. Imposing the cutoff t-value of three raises the number of insignificance to 380 (85%). Even for the 161 significant anomalies, their magnitudes are often much lower than originally reported. Out of the 161, the q-factor model leaves 115 alphas insignificant (150 with t < 3). In all, capital markets are more efficient than previously recognized.  


Friday, April 14, 2017

On Pseudo Out-of-Sample Model Selection

Great to see that Hirano and Wright (HW), "Forecasting with Model Uncertainty", finally came out in Econometrica. (Ungated working paper version here.)

HW make two key contributions. First, they characterize rigorously the source of the inefficiency in forecast model selection by pseudo out-of-sample methods (expanding-sample, split-sample, ...), adding invaluable precision to more intuitive discussions like Diebold (2015). (Ungated working paper version here.) Second, and very constructively, they show that certain simulation-based estimators (including bagging) can considerably reduce, if not completely eliminate, the inefficiency.


Abstract: We consider forecasting with uncertainty about the choice of predictor variables. The researcher wants to select a model, estimate the parameters, and use the parameter estimates for forecasting. We investigate the distributional properties of a number of different schemes for model choice and parameter estimation, including: in‐sample model selection using the Akaike information criterion; out‐of‐sample model selection; and splitting the data into subsamples for model selection and parameter estimation. Using a weak‐predictor local asymptotic scheme, we provide a representation result that facilitates comparison of the distributional properties of the procedures and their associated forecast risks. This representation isolates the source of inefficiency in some of these procedures. We develop a simulation procedure that improves the accuracy of the out‐of‐sample and split‐sample methods uniformly over the local parameter space. We also examine how bootstrap aggregation (bagging) affects the local asymptotic risk of the estimators and their associated forecasts. Numerically, we find that for many values of the local parameter, the out‐of‐sample and split‐sample schemes perform poorly if implemented in the conventional way. But they perform well, if implemented in conjunction with our risk‐reduction method or bagging.

Sunday, June 26, 2016

Regularization for Long Memory

Two earlier regularization posts focused on panel data and generic time series contexts. Now consider a specific time-series context: long memory. For exposition consider the simplest case of a pure long memory DGP,  \( (1-L)^d y_t = \varepsilon_t \) with  \( |d| < 1/2  \).  This \( ARFIMA(0,d,0) \) process is  is \( AR(\infty) \) with very slowly decaying coefficients due to the long memory. If you KNEW the world was was \(ARFIMA(0,d,0)\) you'd just fit \(d\) using GPH or Whittle or whatever, but you're not sure, so you'd like to stay flexible and fit a very long \(AR\) (an \(AR(100) \), say). But such a profligate parameterization is infeasible or at least very wasteful. A solution is to fit the \(AR(100) \) but regularize by estimating with ridge or a LASSO variant, say.

Related, recall the Corsi "HAR" approximation to long memory. It's just a long autoregression subject to coefficient restrictions. So you could do a LASSO estimation, as in Audrino and Knaus (2013). Related analysis and references are in a Humboldt University 2015 master's thesis.)


Finally, note that in all of the above it might be desirable to change the LASSO centering point for shrinage/selection to match the long-memory restriction. (In standard LASSO it's just 0.)

Sunday, February 14, 2016

David Hendry on Measurement, Theory, and Model Selection

David Hendry has an interesting new paper on the interplay of measurement and theory in model selection, "Deciding Between Alternative Approaches in Macroeconomics." As usual, David frames deep and important questions, and he furnishes penetrating insights as he works toward answers. Of course there's lots to disagree with, but that's beside the point -- original thinking is usually unsettling and controversial.

On David's theme of doing credible empirical work while maintaining focus on theory, see the No Hesitations post, "Shrinking VAR's Toward Theory." More generally, on his broad theme of the interplay between measurement and theory, see my The Known, the Unknown and the Unknowable (KuU) in Financial Risk Management, with Neil Doherty and Dick Herring. The introduction is freely available here. And don't forget another earlier No Hesitations post, "Theory gets too Much Respect, and Measurement Doesn't get Enough". (!)