Below are the slides from my discussion of Helene Rey et al., "Answering the Queen: Machine Learning and Financial Crises", which I gave a few days ago at a fine NBER IFM meeting (program and clickable papers here). I also discussed it in June at the BIS annual research meeting in Zurich. The key development since the earlier mid-summer draft is that they actually implemented a real-time financial crisis prediction analysis for France using vintage data, as opposed to quasi-real-time using final-revised data. Moving to real time of course somewhat degrades the quasi-real-time results, but they largely hold up. Very impressive. Therefore I now offer suggestions for improving evaluation credibility in the remaining cases where vintage datasets are not yet available. On the other hand, I also note how subtle but important look-ahead biases can creep in even when vintage data are available and used. I conclude that the only fully-convincing evaluation involves implementing their approach moving forward, recording the results, and building up a true track record.
Econometrics, economics, finance, random rants.
Econometrics, economics, finance, random rants...
Showing posts with label Machine learning. Show all posts
Showing posts with label Machine learning. Show all posts
Sunday, October 27, 2019
Online Learning vs. TVP Forecast Combination
[This post is based on the first slide (below) of a discussion of Helene Rey et al., which I gave a few days ago at a fine NBER IFM meeting (program and clickable papers here). The paper is fascinating and impressive, and I'll blog on it separately next time. But the slide below is more of a side rant on general issues, and I skipped it in the discussion of Rey et al. to be sure to have time to address their particular issues.]
Quite a while ago I blogged here on the ex ante expected loss minimization that underlies traditional econometric/statistical forecast combination, vs. the ex post regret minimization that underlies "online learning" and related "machine learning" methods. Nothing has changed. That is, as regards ex post regret minimization, I'm still intrigued, but I'm still not persuaded.
And there's another thing that bothers me. As implemented, ML-style online learning and traditional econometric-style forecast combination with time-varying parameters (TVPs) are almost identical: just projection (regression) of realizations on forecasts, reading off the combining weights as the regression coefficients. OF COURSE we can generalize to allow for time-varying combining weights, non-linear combinations, regularization in high dimensions, etc., and hundreds of econometrics papers have addressed and explored those issues. Yet the ML types seem to think they invented everything, and too many economists are buying it. Rey et al., for example, don't so much as mention the econometric forecast combination literature, which by now occupies large chapters of leading textbooks, like Elliott and Timmermann at the bottom of the slide below.
Quite a while ago I blogged here on the ex ante expected loss minimization that underlies traditional econometric/statistical forecast combination, vs. the ex post regret minimization that underlies "online learning" and related "machine learning" methods. Nothing has changed. That is, as regards ex post regret minimization, I'm still intrigued, but I'm still not persuaded.
And there's another thing that bothers me. As implemented, ML-style online learning and traditional econometric-style forecast combination with time-varying parameters (TVPs) are almost identical: just projection (regression) of realizations on forecasts, reading off the combining weights as the regression coefficients. OF COURSE we can generalize to allow for time-varying combining weights, non-linear combinations, regularization in high dimensions, etc., and hundreds of econometrics papers have addressed and explored those issues. Yet the ML types seem to think they invented everything, and too many economists are buying it. Rey et al., for example, don't so much as mention the econometric forecast combination literature, which by now occupies large chapters of leading textbooks, like Elliott and Timmermann at the bottom of the slide below.
Saturday, March 23, 2019
Big Data in Dynamic Predictive Modeling
Our Journal of Econometrics issue, Big Data in Dynamic Predictive Econometric Modeling, is now in press. It is partly based on a Penn conference, generously supported by Penn's Warren Center for Network and Data Sciences, University of Chicago's Stevanovich Center for Financial Mathematics, and Penn's Institute for Economic Research. The intro is here and the paper list is here.
Friday, March 15, 2019
Neyman-Pearson Classification
Neyman-Pearson (NP) hypothesis testing insists on fixed asymptotic test size (5%, say) and then takes whatever power it can get. Bayesian hypothesis assessment, in contrast, treats type I and II errors symmetrically, with size approaching 0 and power approaching 1 asymptotically.
Classification tends to parallel Bayesian hypothesis assessment, again treating type I and II errors symmetrically. For example, I might do a logit regression and classify cases with fitted P(I=1)<1/2 as group 0 and cases with fitted P(I=1)>1/2 as group 1. The classification threshold of 1/2 produces a ``Bayes classifier".
Bayes classifiers seem natural, and in many applications they are. But an interesting insight is that some classification problems may have hugely different costs of type I and II errors, in which case an NP classification approach may be entirely natural, not clumsy. (Consider, for example, deciding whether to convict someone of a crime that carries the death penalty. Many people would view the cost of a false declaration of "guilty" as much greater than the cost of a false "innocent".)
Classification tends to parallel Bayesian hypothesis assessment, again treating type I and II errors symmetrically. For example, I might do a logit regression and classify cases with fitted P(I=1)<1/2 as group 0 and cases with fitted P(I=1)>1/2 as group 1. The classification threshold of 1/2 produces a ``Bayes classifier".
Bayes classifiers seem natural, and in many applications they are. But an interesting insight is that some classification problems may have hugely different costs of type I and II errors, in which case an NP classification approach may be entirely natural, not clumsy. (Consider, for example, deciding whether to convict someone of a crime that carries the death penalty. Many people would view the cost of a false declaration of "guilty" as much greater than the cost of a false "innocent".)
This leads to the idea and desirability of NP classifiers. The issue is how to bound the type I classification error probability at some small chosen value. Obviously it involves moving the classification threshold away from 1/2, but figuring out exactly what to do turns out to be a challenging problem. Xin Tong and co-authors have made good progress. Here are some of his papers (from his USC site):
- Chen, Y., Li, J.J., and Tong, X.* (2019) Neyman-Pearson criterion (NPC): a model selection criterion for asymmetric binary classification. arXiv:1903.05262.
- Tong, X., Xia, L., Wang, J., and Feng, Y. (2018) Neyman-Pearson classification: parametrics and power enhancement. arXiv:1802.02557v3.
- Xia, L., Zhao, R., Wu, Y., and Tong, X.* (2018) Intentional control of type I error over unconscious data distortion: a Neyman-Pearson approach to text classification. arXiv:1802.02558.
- Tong, X.*, Feng, Y. and Li, J.J. (2018) Neyman-Pearson (NP) classification algorithms and NP receiver operating characteristics (NP-ROC). Science Advances, 4(2):eaao1659.
- Zhao, A., Feng, Y., Wang, L., and Tong, X.* (2016) Neyman-Pearson classification under high-dimensional settings. Journal of Machine Learning Research, 17:1−39.
- Li, J.J. and Tong, X. (2016) Genomic applications of the Neyman-Pearson classification paradigm. Chapter in Big Data Analytics in Genomics. Springer (New York). DOI: 10.1007/978-3-319-41279-5; eBook ISBN: 978-3-319-41279-5.
- Tong, X.*, Feng, Y. and Zhao, A. (2016) A survey on Neyman-Pearson classification and suggestions for future research. Wiley Interdisciplinary Reviews: Computational Statistics, 8:64-81.
- Tong, X.* (2013). A plug-in approach to Neyman-Pearson classification. Journal of Machine Learning Research, 14:3011-3040.
- Rigollet, P. and Tong, X. (2011) Neyman-Pearson classification, convexity and stochastic constraints. Journal of Machine Learning Research, 12:2825-2849.
Friday, January 25, 2019
Network Data and Machine Learning
This just arrived, announcing an upcoming conference on the ML/networks interface. It's definitely worth reading through the synopsis and topics and titles and authors.
"An exciting workshop on Machine Learning for Network Data is taking place at New York University on January 29. The event will discuss emerging challenges on generalizing the successes of image and speech processing to information domains with irregular structure. The workshop includes highlight talks by Yann LeCun and Brian Sadler as well as short talks by a collection of national leaders in the development of machine learning techniques for processing network data. The event is free to attend and open to the public but registration is required because of space limitations. Please visit the workshop site to access the registration form."
Monday, January 21, 2019
Machine Learning for Economists
My Penn colleague Jesus Fernandez-Villaverde has a nice slide deck here. He asked me to warn you that this is a highly-preliminary version (0.1!), and to thank, without implicating, Stephen Hansen, as the deck draws on joint work.
Friday, September 14, 2018
Machine Learning for Forecast Combination
How could I have forgotten to announce my latest paper, "Machine Learning for Regularized Survey Forecast Combination: Partially-Egalitarian Lasso and its Derivatives"? (Actually a heavily-revised version of an earlier paper, including a new title.) Came out as an NBER w.p. a week or two ago.
Tuesday, July 24, 2018
Gu-Kelly-Xiu and Neural Nets in Economics
I'm on record as being largely unimpressed by the contributions of neural nets (NN's) in economics thus far. In many economic environments the relevant non-linearities seem too weak and the signal/noise ratios too low for NN's to contribute much.
The Gu-Kelly-Xiu paper that I mentioned earlier may change that. I mentioned their success in applying machine-learning methods to forecast equity risk premia out of sample. NN's, in particular, really shine. The paper is thoroughly and meticulously done.
This is potentially a really big deal.
The Gu-Kelly-Xiu paper that I mentioned earlier may change that. I mentioned their success in applying machine-learning methods to forecast equity risk premia out of sample. NN's, in particular, really shine. The paper is thoroughly and meticulously done.
This is potentially a really big deal.
Thursday, July 19, 2018
Machine Learning, Volatility, and the Interface
Just got back from the NBER Summer Institute. Lots of good stuff happening in the Forecasting and Empirical Methods group. The program, with links to papers, is here.
But there's actually a big interface.
Lots of room for extensions too. Here's a great example. Consider the interface of the Gu-Kelly-Xiu and Bollerslev-Patton-Quagvleg papers. At first you might think that there is no interface.
Kelly-Xiu is about using off-the-shelf machine-learning methods to model risk premia in financial markets; that is, to construct portfolios that deliver superior performance. (I had guessed they'd get nothing, but I was massively wrong.) Bollerslev et al. is about predicting realized covariance by exploiting info on past signs (e.g., was yesterday's covariance cross-product pos-pos, neg-neg, pos-neg, or neg-pos?). (They also get tremendous results.)
Kelly-Xiu is about using off-the-shelf machine-learning methods to model risk premia in financial markets; that is, to construct portfolios that deliver superior performance. (I had guessed they'd get nothing, but I was massively wrong.) Bollerslev et al. is about predicting realized covariance by exploiting info on past signs (e.g., was yesterday's covariance cross-product pos-pos, neg-neg, pos-neg, or neg-pos?). (They also get tremendous results.)
But there's actually a big interface.
Note that Kelly-Xiu is about conditional mean dynamics -- uncovering the determinants of expected excess returns. You might expect even better results for derivative assets, as the volatility dynamics that drive options prices may be nonlinear in ways missed by standard volatility models. And that's exactly the flavor of the Bollerslev et al. results -- they find that a tree structure conditioning on sign is massively successful.
But Bollerslev et al. don't do any machine learning. Instead they basically stumble upon their result, guided by their fine intuition. So here's a fascinating issue to explore: Hit the Bollerslev et al. realized covariance data with machine learning (in particular, tree methods like random forests) and see what happens. Does it "discover" the Bollerslev et al. result? If not, why not, and what does it discover? Does it improve upon Bollerslev et al.?
Thursday, June 7, 2018
Machines Learning Finance
FRB Atlanta recently hosted a meeting on "Machines Learning Finance". Kind of an ominous, threatening (Orwellian?) title, but there were lots of (non-threatening...) pieces. I found the surveys by Ryan Adams and John Cunningham particularly entertaining. A clear theme on display throughout the meeting was that "supervised learning" -- the main strand of machine learning -- is just function estimation, and in particular, conditional mean estimation. That is, regression. It may involve high dimensions, non-linearities, binary variables, etc., but at the end of the day it's still just regression. If you're a regular No Hesitations reader, the "insight" that supervised learning = regression will hardly be novel to you, but still it's good to see it disseminating widely.
Monday, April 2, 2018
Econometrics, Machine Learning, and Big Data
Here's a useful slide deck by Greg Duncan at Amazon, from a recent seminar at FRB San Francisco (powerpoint, ughhh, sorry...). It's basically a superset of the keynote talk he gave at Penn's summer 2017 conference, Big Data in Predictive Dynamic Econometric Modeling. Greg understands better than most the close connection between "machine learning" and econometrics / statistics, especially between machine learning and the predictive perspective emphasized in time series for a century or so.
Monday, February 26, 2018
STILL MORE on NN's and ML
I recently discussed how the nonparametric consistency of wide NN's proved underwhelming, which is partly why econometricians lost interest in NN's in the 1990s.
The other thing was the realization that NN objective surfaces are notoriously bumpy, so that arrival at a local optimum (e.g., by the stochastic gradient descent popular in NN circles) offered little comfort.
So econometricians' interest declined on both counts. But now both issues are being addressed. The new focus on NN depth as opposed to width is bearing much fruit. And recent advances in "reinforcement learning" methods effectively promote global as opposed to just local optimization, by experimenting (injecting randomness) in clever ways. (See, e.g., Taddy section 6, here.)
All told, it seems like quite an exciting new time for NN's. I've been away for 15 years. Time to start following again...
The other thing was the realization that NN objective surfaces are notoriously bumpy, so that arrival at a local optimum (e.g., by the stochastic gradient descent popular in NN circles) offered little comfort.
So econometricians' interest declined on both counts. But now both issues are being addressed. The new focus on NN depth as opposed to width is bearing much fruit. And recent advances in "reinforcement learning" methods effectively promote global as opposed to just local optimization, by experimenting (injecting randomness) in clever ways. (See, e.g., Taddy section 6, here.)
All told, it seems like quite an exciting new time for NN's. I've been away for 15 years. Time to start following again...
Monday, February 19, 2018
More on Neural Nets and ML
I earlier mentioned Matt Taddy's "The Technological Elements of Artificial Intelligence" (ungated version here).
Among other things the paper has good perspective on the past and present of neural nets. (Read: his views mostly, if not exactly, match mine...)
Here's my personal take on some of the history vis a vis econometrics:
Econometricians lost interest in NN's in the 1990's. The celebrated Hal White et al. proof of NN non-parametric consistency as NN width (number of neurons) gets large at an appropriate rate was ultimately underwhelming, insofar as it merely established for NN's what had been known for decades for various other non-parametric estimators (kernel, series, nearest-neighbor, trees, spline, etc.). That is, it seemed that there was nothing special about NN's, so why bother?
But the non-parametric consistency focus was all on NN width; no one thought or cared much about NN depth. Then, more recently, people noticed that adding NN depth (more hidden layers) could be seriously helpful, and the "deep learning" boom took off.
Here are some questions/observations on the new "deep learning":
1. Adding NN depth often seems helpful, insofar as deep learning often seems to "work" in various engineering applications, but where/what are the theorems? What can be said rigorously about depth?
2. Taddy emphasizes what might be called two-step deep learning. In the first step, "pre-trained" hidden layer nodes are obtained based on unsupervised learning (e.g., principle components (PC)) from various sets of variables. And then the second step proceeds as usual. That's very similar to the age-old idea of PC regression. Or, in multivariate dynamic environments and econometrics language, "factor-augmented vector autoregression" (FAVAR), as in Bernanke et al. (2005). So, are modern implementations of deep NN's effectively just nonlinear FAVAR's? If so, doesn't that also seem underwhelming, in the sense of -- dare I say it -- there being nothing really new about deep NN's?
3. Moreover, PC regressions and FAVAR's have issues of their own relative to one-step procedures like ridge or LASSO. See this and this.
Among other things the paper has good perspective on the past and present of neural nets. (Read: his views mostly, if not exactly, match mine...)
Here's my personal take on some of the history vis a vis econometrics:
Econometricians lost interest in NN's in the 1990's. The celebrated Hal White et al. proof of NN non-parametric consistency as NN width (number of neurons) gets large at an appropriate rate was ultimately underwhelming, insofar as it merely established for NN's what had been known for decades for various other non-parametric estimators (kernel, series, nearest-neighbor, trees, spline, etc.). That is, it seemed that there was nothing special about NN's, so why bother?
But the non-parametric consistency focus was all on NN width; no one thought or cared much about NN depth. Then, more recently, people noticed that adding NN depth (more hidden layers) could be seriously helpful, and the "deep learning" boom took off.
Here are some questions/observations on the new "deep learning":
1. Adding NN depth often seems helpful, insofar as deep learning often seems to "work" in various engineering applications, but where/what are the theorems? What can be said rigorously about depth?
2. Taddy emphasizes what might be called two-step deep learning. In the first step, "pre-trained" hidden layer nodes are obtained based on unsupervised learning (e.g., principle components (PC)) from various sets of variables. And then the second step proceeds as usual. That's very similar to the age-old idea of PC regression. Or, in multivariate dynamic environments and econometrics language, "factor-augmented vector autoregression" (FAVAR), as in Bernanke et al. (2005). So, are modern implementations of deep NN's effectively just nonlinear FAVAR's? If so, doesn't that also seem underwhelming, in the sense of -- dare I say it -- there being nothing really new about deep NN's?
3. Moreover, PC regressions and FAVAR's have issues of their own relative to one-step procedures like ridge or LASSO. See this and this.
Tuesday, February 13, 2018
Neural Nets, ML and AI
"The Technological Elements of Artificial Intelligence", by Matt Taddy, is packed with insight on the development of neural nets and ML as related to the broader development of AI. I have lots to say, but it will have to wait until next week. For now I just want you to have the paper. Ungated version at http://www.nber.org/chapters/c14021.pdf.
Abstract:
We have seen in the past decade a sharp increase in the extent that companies use data to optimize their businesses. Variously called the `Big Data' or `Data Science' revolution, this has been characterized by massive amounts of data, including unstructured and nontraditional data like text and images, and the use of fast and flexible Machine Learning (ML) algorithms in analysis. With recent improvements in Deep Neural Networks (DNNs) and related methods, application of high-performance ML algorithms has become more automatic and robust to different data scenarios. That has led to the rapid rise of an Artificial Intelligence (AI) that works by combining many ML algorithms together - each targeting a straightforward prediction task - to solve complex problems.
We will define a framework for thinking about the ingredients of this new ML-driven AI. Having an understanding of the pieces that make up these systems and how they fit together is important for those who will be building businesses around this technology. Those studying the economics of AI can use these definitions to remove ambiguity from the conversation on AI's projected productivity impacts and data requirements. Finally, this framework should help clarify the role for AI in the practice of modern business analytics and economic measurement.
We have seen in the past decade a sharp increase in the extent that companies use data to optimize their businesses. Variously called the `Big Data' or `Data Science' revolution, this has been characterized by massive amounts of data, including unstructured and nontraditional data like text and images, and the use of fast and flexible Machine Learning (ML) algorithms in analysis. With recent improvements in Deep Neural Networks (DNNs) and related methods, application of high-performance ML algorithms has become more automatic and robust to different data scenarios. That has led to the rapid rise of an Artificial Intelligence (AI) that works by combining many ML algorithms together - each targeting a straightforward prediction task - to solve complex problems.
We will define a framework for thinking about the ingredients of this new ML-driven AI. Having an understanding of the pieces that make up these systems and how they fit together is important for those who will be building businesses around this technology. Those studying the economics of AI can use these definitions to remove ambiguity from the conversation on AI's projected productivity impacts and data requirements. Finally, this framework should help clarify the role for AI in the practice of modern business analytics and economic measurement.
Monday, February 12, 2018
ML, Forecasting, and Market Design
Nice stuff from Milgrom and Tadelis. Improved forecasting via improved machine learning in turn helps improve our ability to design effective markets -- better anticipating consumer/producer demand/supply movements, more finely segmenting and targeting consumers/producers, more accurately setting auction reserve prices, etc. Presumably full density forecasts, not just the point forecasts on which ML tends to focus, should soon move to center stage.
http://www.nber.org/chapters/c14008.pdf
http://www.nber.org/chapters/c14008.pdf
Monday, February 5, 2018
Big Data, Machine Learning, and Economic Statistics
Greetings from a very happy Philadelphia celebrating the Eagles' victory!
The following is adapted from the "background" and "purpose" statements for a planned 2019 NBER/CRIW conference, "Big Data for 21st Century Economic Statistics". Prescient and fascinating reading. (The full call for papers is here.)
Background: The coming decades will witness significant changes in the production of the social and economic statistics on which government officials, business decision makers, and private citizens rely. The statistical information currently produced by the federal statistical agencies rests primarily on “designed data” -- that is, data collected through household and business surveys. The increasing cost of fielding these surveys, the difficulty of obtaining survey responses, and questions about the reliability of some of the information collected, have raised questions about the sustainability of that model. At the same time, the potential for using “big data” -- very large data sets built to meet governments’ and businesses’ administrative and operational needs rather than for statistical purposes -- in the production of official statistics has grown.
These naturally-occurring data include not only administrative data maintained by government agencies but also scanner data, data scraped from the Web, credit card company records, data maintained by payroll providers, medical records, insurance company records, sensor data, and the Internet of Things. If the challenges associated with their use can be satisfactorily resolved, these emerging sorts of data could allow the statistical agencies not only to supplement or replace the survey data on which they currently depend, but also to introduce new statistics that are more granular, more up-to-date, and of higher quality than those currently being produced.
Purpose: The purpose of this conference is to provide a forum where economists, data providers, and data analysts can meet to present research on the use of big data in the production of federal social and economic statistics. Among other things, this involves discussing (1) Methods for combining multiple data sources, whether they be carefully designed surveys or experiments, large government administrative datasets, or private sector big data, to produce economic and social statistics; (2) Case studies illustrating how big data can be used to improve or replace existing statistical data series or create new statistical data series; (3) Best practices for characterizing the quality of big data sources and blended estimates constructed using data from multiple sources.
The following is adapted from the "background" and "purpose" statements for a planned 2019 NBER/CRIW conference, "Big Data for 21st Century Economic Statistics". Prescient and fascinating reading. (The full call for papers is here.)
Background: The coming decades will witness significant changes in the production of the social and economic statistics on which government officials, business decision makers, and private citizens rely. The statistical information currently produced by the federal statistical agencies rests primarily on “designed data” -- that is, data collected through household and business surveys. The increasing cost of fielding these surveys, the difficulty of obtaining survey responses, and questions about the reliability of some of the information collected, have raised questions about the sustainability of that model. At the same time, the potential for using “big data” -- very large data sets built to meet governments’ and businesses’ administrative and operational needs rather than for statistical purposes -- in the production of official statistics has grown.
These naturally-occurring data include not only administrative data maintained by government agencies but also scanner data, data scraped from the Web, credit card company records, data maintained by payroll providers, medical records, insurance company records, sensor data, and the Internet of Things. If the challenges associated with their use can be satisfactorily resolved, these emerging sorts of data could allow the statistical agencies not only to supplement or replace the survey data on which they currently depend, but also to introduce new statistics that are more granular, more up-to-date, and of higher quality than those currently being produced.
Purpose: The purpose of this conference is to provide a forum where economists, data providers, and data analysts can meet to present research on the use of big data in the production of federal social and economic statistics. Among other things, this involves discussing (1) Methods for combining multiple data sources, whether they be carefully designed surveys or experiments, large government administrative datasets, or private sector big data, to produce economic and social statistics; (2) Case studies illustrating how big data can be used to improve or replace existing statistical data series or create new statistical data series; (3) Best practices for characterizing the quality of big data sources and blended estimates constructed using data from multiple sources.
Wednesday, November 8, 2017
Artificial Intelligence, Machine Learning, and Productivity
As Bob Solow famously quipped, "You can see the computer age everywhere but in the productivity statistics". That was in 1987. The new "Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics," NBER w.p. 24001, by Brynjolfsson, Rock, and Syverson, brings us up to 2017. Still a puzzle. Fascinating. Ungated version here.
Sunday, October 22, 2017
Pockets of Predictability
The possibility of localized "pockets of predictability", particularly in financial markets, is obviously intriguing. Recently I'm noticing a similarly-intriguing pocket of research on pockets of predictability.
The following paper, for example, was presented at 2017 the NBER-NSF Time Series conference at Northwestern University, even if it is evidently not yet circulating:
The following paper, for example, was presented at 2017 the NBER-NSF Time Series conference at Northwestern University, even if it is evidently not yet circulating:
"Pockets of Predictability", by Leland Farmer (UCSD), Lawrence Schmidt (Chicago), and Allan Timmermann (UCSD). Abstract: We show that return predictability in the U.S. stock market is a localized phenomenon, in which short periods, “pockets,” with significant predictability are interspersed with long periods with little or no evidence of return predictability. We explore possible explanations of this finding, including time-varying risk premia, and find that they are inconsistent with a general class of affine asset pricing models which allow for stochastic volatility and compound Poisson jumps. We find that pockets of return predictability can, however, be explained by a model of incomplete learning in which the underlying cash flow process is subject to change and investors update their priors about the current state. Simulations from the model demonstrate that investors’ learning about the underlying cash flow process can induce patterns that look, ex-post, like local return predictability, even in a model in which ex-ante expected returns are constant.
And this one just appeared as an NBER w.p.: "Sparse Signals in the Cross-Section of Returns", by Alexander M. Chinco, Adam D. Clark-Joseph, Mao Ye, NBER w.p. 23933, October 2017.
http://papers.nber.org/papers/w23933?utm_campaign=ntw&utm_medium=email&utm_source=ntw
Abstract: This paper applies the Least Absolute Shrinkage and Selection Operator (LASSO) to make rolling 1-minute-ahead return forecasts using the entire cross section of lagged returns as candidate predictors. The LASSO increases both out-of-sample fit and forecast-implied Sharpe ratios. And, this out-of-sample success comes from identifying predictors that are unexpected, short-lived, and sparse. Although the LASSO uses a statistical rule rather than economic intuition to identify predictors, the predictors it identifies are nevertheless associated with economically meaningful events: the LASSO tends to identify as predictors stocks with news about fundamentals.
Here's some associated work in dynamical systems theory: "A Mechanism for Pockets of Predictability in Complex Adaptive Systems", by Jorgen Vitting Andersen, Didier Sornette, Europhysics Letters, 2005. https://arxiv.org/abs/cond-mat/0410762
Abstract: We document a mechanism operating in complex adaptive systems leading to dynamical pockets of predictability ("prediction days''), in which agents collectively take predetermined courses of action, transiently decoupled from past history. We demonstrate and test it out-of-sample on synthetic minority and majority games as well as on real financial time series. The surprising large frequency of these prediction days implies a collective organization of agents and of their strategies which condense into transitional herding regimes.
There's even an ETH Zürich master's thesis: "In Search Of Pockets Of Predictability", by AT Morera, 2008
https://www.ethz.ch/content/dam/ethz/special-interest/mtec/chair-of-entrepreneurial-risks-dam/documents/dissertation/master%20thesis/Master_Thesis_Alan_Taxonera_Sept08.pdf
Finally, related ideas have appeared recently in the forecast evaluation literature, such as this paper and many of the references therein: "Testing for State-Dependent Predictive Ability", by Sebastian Fossati, University of Alberta, September 2017.Abstract: We document a mechanism operating in complex adaptive systems leading to dynamical pockets of predictability ("prediction days''), in which agents collectively take predetermined courses of action, transiently decoupled from past history. We demonstrate and test it out-of-sample on synthetic minority and majority games as well as on real financial time series. The surprising large frequency of these prediction days implies a collective organization of agents and of their strategies which condense into transitional herding regimes.
There's even an ETH Zürich master's thesis: "In Search Of Pockets Of Predictability", by AT Morera, 2008
https://www.ethz.ch/content/dam/ethz/special-interest/mtec/chair-of-entrepreneurial-risks-dam/documents/dissertation/master%20thesis/Master_Thesis_Alan_Taxonera_Sept08.pdf
https://sites.ualberta.ca/~econwps/2017/wp2017-09.pdf
Abstract: This paper proposes a new test for comparing the out-of-sample forecasting performance of two competing models for situations in which the predictive content may be state-dependent (for example, expansion and recession states or low and high volatility states). To apply this test the econometrician is not required to observe when the underlying states shift. The test is simple to implement and accommodates several different cases of interest. An out-of-sample forecasting exercise for US output growth using real-time data illustrates the improvement of this test over previous approaches to perform forecast comparison.
Saturday, October 14, 2017
Machine Learning and Macro
Earlier I posted here on machine learning and central banking. Here's something related.
Last week Penn's Warren Center hosted a timely and stimulating conference, "Machine Learning for Macroeconomic Prediction and Policy". The program appears below. Papers were not posted, but with a little Googling you should be able to obtain those that are available.
Conference on Machine Learning for Macroeconomic Prediction and Policy
October 12 and 13, 2017
Glandt Forum, Singh Center for Nanotechnology
Co-Sponsored by Penn’s Warren Center for Network and Data Sciences
and the Federal Reserve Bank of Philadelphia
Organizers: Michael Dotsey (FRBP), Jesus Fernandez-Villaverde (Penn), Michael Kearns (Penn)
SCHEDULE:
Thursday October 12:
8:00 Breakfast
8:45 Welcome
9:00 Stephen Hansen (University of Oxford): The Long-Run Information Effect of Central Bank Text
9:45 Stephen Ryan (Washington University): Classi cation Trees for Heterogeneous Moment-Based Models
10:30 Break
11:00 James Cowie (DeepMacro): DeepMacro Data Challenges
11:45 Galo Nuno (Banco de España): Machine Learning and Heterogeneous Agent Models
12:30 Lunch
1:30: Francis X. Diebold (Penn): Egalitarian LASSO for Combining Central Bank Survey Forecasts
2:15 Lyle Ungar (Penn): How to Make Better Forecasts
3:00 Vegard Larsen (Norges Bank): Components of Uncertainty
3:45 Break
4:15 Panel: ML and Econometrics: Similarities and Differences (Michael Kearns, Vegard Larsen, Stephen Hansen, Rakesh Vohra (Penn))
Friday October 13:
9:00 Aaron Smalter Hall (Federal Reserve Bank of Kansas City): Recession Forecasting with Bayesian Classification
9: 45 Susan Athey (Stanford GSB): Estimating Heterogeneity in Structural Parameters Using Generalized Random Forests
10:30 Break
11:00 Panel: ML Challenges at the Fed (Jose Canals-Cerda (Philadelphia Fed), Galo Nuno, Jesus Fernandez-Villaverde, Aaron Smalter Hall)
12:30 Lunch
Departures
Last week Penn's Warren Center hosted a timely and stimulating conference, "Machine Learning for Macroeconomic Prediction and Policy". The program appears below. Papers were not posted, but with a little Googling you should be able to obtain those that are available.
Conference on Machine Learning for Macroeconomic Prediction and Policy
October 12 and 13, 2017
Glandt Forum, Singh Center for Nanotechnology
Co-Sponsored by Penn’s Warren Center for Network and Data Sciences
and the Federal Reserve Bank of Philadelphia
Organizers: Michael Dotsey (FRBP), Jesus Fernandez-Villaverde (Penn), Michael Kearns (Penn)
SCHEDULE:
Thursday October 12:
8:00 Breakfast
8:45 Welcome
9:00 Stephen Hansen (University of Oxford): The Long-Run Information Effect of Central Bank Text
9:45 Stephen Ryan (Washington University): Classi cation Trees for Heterogeneous Moment-Based Models
10:30 Break
11:00 James Cowie (DeepMacro): DeepMacro Data Challenges
11:45 Galo Nuno (Banco de España): Machine Learning and Heterogeneous Agent Models
12:30 Lunch
1:30: Francis X. Diebold (Penn): Egalitarian LASSO for Combining Central Bank Survey Forecasts
2:15 Lyle Ungar (Penn): How to Make Better Forecasts
3:00 Vegard Larsen (Norges Bank): Components of Uncertainty
3:45 Break
4:15 Panel: ML and Econometrics: Similarities and Differences (Michael Kearns, Vegard Larsen, Stephen Hansen, Rakesh Vohra (Penn))
Friday October 13:
9:00 Aaron Smalter Hall (Federal Reserve Bank of Kansas City): Recession Forecasting with Bayesian Classification
9: 45 Susan Athey (Stanford GSB): Estimating Heterogeneity in Structural Parameters Using Generalized Random Forests
10:30 Break
11:00 Panel: ML Challenges at the Fed (Jose Canals-Cerda (Philadelphia Fed), Galo Nuno, Jesus Fernandez-Villaverde, Aaron Smalter Hall)
12:30 Lunch
Departures
Sunday, September 24, 2017
Egalitarian LASSO for Forecast Combination
Here's a new one. It was something of a long and winding road. We introduce simple "egalitarian LASSO" procedures that set some combining weights to zero and shrink those remaining toward equality. The feasible versions don't work very well, due do difficulties associated with cross-validating tuning parameters in small samples, but the lessons learned in studying the infeasible version turn out to be very valuable -- indeed they directly motivate a new procedure, which we call "best <N-averaging", which solves the cross-validation problem and performs intriguingly well.
Diebold, F.X. and Shin, M. (2017), “Beating the Simple Average: Egalitarian LASSO for Combining Economic Forecasts”, Penn Institute for Economic Research (PIER) Working Paper No. 17-017, available at SSRN: https://ssrn.com/abstract=3032492.
Diebold, F.X. and Shin, M. (2017), “Beating the Simple Average: Egalitarian LASSO for Combining Economic Forecasts”, Penn Institute for Economic Research (PIER) Working Paper No. 17-017, available at SSRN: https://ssrn.com/abstract=3032492.
Subscribe to:
Posts (Atom)








