Econometrics, economics, finance, random rants.

Econometrics, economics, finance, random rants...
Showing posts with label Factor Structure. Show all posts
Showing posts with label Factor Structure. Show all posts

Saturday, October 19, 2019

Missing data in Factor Models

Serena Ng's site is back. Her new paper with Jushan Bai, "Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data", which I blogged about earlier, is now up on arXiv, here.  Closely related, Marcus Pelger just sent me his new paper with Ruoxuan Xiong, "Large Dimensional Latent Factor Modeling with Missing Observations and Applications to Causal Inference", which I look forward to reading. One is dated Oct 15 and one is dated Oct 16. Science is rushing forward!

Sunday, October 6, 2019

Large Dimensional Factor Analysis with Missing Data

Back from the very strong Stevanovich meeting.  Program and abstracts here.  One among many  highlights was:

Large  Dimensional  Factor Analysis  with  Missing Data
Presented by Serena Ng, (Columbia, Dept. of Economics)

Abstract:
This paper introduces two factor-based imputation procedures that will fill
missing values with consistent estimates of the common component. The
first method is applicable when the missing data are bunched. The second
method is appropriate when the data are missing in a staggered or
disorganized manner. Under the strong factor assumption, it is shown that
the low rank component can be consistently estimated but there will be at
least four convergence rates, and for some entries, re-estimation can
accelerate convergence. We provide a complete characterization of the
sampling error without requiring regularization or imposing the missing at
random assumption as in the machine learning literature. The
methodology can be used in a wide range of applications, including
estimation of covariances and counterfactuals.

This paper just blew me away.  Re-arrange the X columns to get all the "complete cases across people" (tall block) in the leftmost columns, and re-arrange the X rows to get all the "complete cases across variables" (wide block) in the topmost rows.  The intersection is the "balanced" block in the upper left.  Then iterate on the tall and wide blocks to impute the missing data in the bottom right "missing data" block. The key figure that illustrates the procedure provided a real "eureka moment" for me.  Plus they have a full asymptotic theory as opposed to just worst-case bounds.

Kudos!

I'm not sure whether the paper is circulating yet, and Serena's web site vanished recently (not her fault -- evidently Google made a massive error), but you'll soon be able to get the paper one way or another.






Tuesday, August 7, 2018

Factor Model w Time-Varying Loadings

Markus Pelger has a nice paper on factor modeling with time-varying loadings in high dimensions. There are many possible applications. He applies it to level-slope-curvature yield-curve models. 

For me another really interesting application would be measuring connectedness in financial markets, as a way of tracking systemic risk. The Diebold-Yilmaz (DY) connectedness framework is based on a high-dimensional VAR with time-varying coefficients, but not factor structure. An obvious alternative in financial markets, which we used to discuss a lot but never pursued, is factor structure with time-varying loadings, exactly in Pelger! 

It would seem, however, that any reasonable connectedness measure in a factor environment would need to be based not only time-varying loadings but also time-varying idiosynchratic shock variances, or more precisely a time-varying noise/signal ratio (e.g., in a 1-factor model, the ratio of the idiosyncratic shock variance to the factor innovation variance). That is, connectedness in factor environments is driven by BOTH the size of the loadings on the factor(s) AND the amount of variation in the data explained by the factor(s). Time-varying loadings don't really change anything if the factors are swamped by massive noise. 

Typically one might fix the factor innovation variance for identification, but allow for time-varying idiosyncratic shock variance in addition to time-varying factor loadings. It seems that Pelger's framework does allow for that. Crudely, and continuing the 1-factor example, consider y_t  =  lambda_t  f_t  +  e_t. His methods deliver estimates of the time series of loadings lambda_t and factor f_t, robust to heteroskedasticity in the idiosyncratic shock e_t. Then in a second step one could back out an estimate of the time series of e_t and fit a volatility model to it. 
Then the entire system would be estimated and one could calculate connectedness measures based, for example, on variance decompositions as in the DY framework

Monday, January 8, 2018

Yield-Curve Modeling


Happy New Year to all!

Riccardo Rebonato's Bond Pricing and Yield-Curve Modeling: A Structural Approach will soon appear from Cambridge University Press. It's very well done -- a fine blend of  theory, empirics, market sense, and good prose.  And not least, endearing humility, well-captured by a memorable sentence from the acknowledgements: "My eight-year-old son has forgiven me, I hope, for not playing with him as much as I would have otherwise; perhaps he has been so understanding because he has had a chance to build a few thousand paper planes with the earlier drafts of this book."

TOC below.  Pre-order here

Contents

Acknowledgements page ix
Symbols and Abbreviations xi

Part I The Foundations
1 What This Book Is About 3
2 Definitions, Notation and a Few Mathematical Results 24
3 Links among Models, Monetary Policy and the Macroeconomy 49
4 Bonds: Their Risks and Their Compensations 63
5 The Risk Factors in Action 81
6 Principal Components: Theory 98
7 Principal Components: Empirical Results 108

Part II The Building Blocks: A First Look
8 Expectations 137
9 Convexity: A First Look 147
10 A Preview: A First Look at the Vasicek Model 160

Part III The Conditions of No-Arbitrage
11 No-Arbitrage in Discrete Time 185
12 No-Arbitrage in Continuous Time 196
13 No-Arbitrage with State Price Deflators 206
14 No-Arbitrage Conditions for Real Bonds 224
15 The Links with an Economics-Based Description of Rates 241

Part IV Solving the Models
16 Solving Affine Models: The Vasicek Case 263
17 First Extensions 285
18 A General Pricing Framework 299
19 The Shadow Rate: Dealing with a Near-Zero Lower Bound 329

Part V The Value of Convexity
20 The Value of Convexity 351
21 A Model-Independent Approach to Valuing Convexity 371
22 Convexity: Empirical Results 391

Part VI Excess Returns
23 Excess Returns: Setting the Scene 415
24 Risk Premia, the Market Price of Risk and Expected Excess Returns 431
25 Excess Returns: Empirical Results 449
26 Excess Returns: The Recent Literature – I 473
27 Excess Returns: The Recent Literature – II 497
28 Why Is the Slope a Good Predictor? 527
29 The Spanning Problem Revisited 547

Part VII What the Models Tell Us
30 The Doubly Mean-Reverting Vasicek Model 559
31 Real Yields, Nominal Yields and Inflation: The D’Amico–Kim–Wei Model 575
32 From Snapshots to Structural Models: The Diebold–Rudebusch Approach 602
33 Principal Components as State Variables of Affine Models: The PCA Affine Approach 618
34 Generalizations: The Adrian–Crump–Moench Model 663
35 An Affine, Stochastic-Market-Price-of-Risk Model 688

36 Conclusions 714

Bibliography 725

index 000

Sunday, November 5, 2017

Regression on Term Structures

An important insight regarding use of dynamic Nelson Siegel (DNS) and related term-structure modeling strategies (see here and here) is that they facilitate regression on an entire term structure.  Regressing something on a curve might initially sound strange, or ill-posed.  The insight, of course, is that DNS distills curves into level, slope, and curvature factors; hence if you know the factors, you know the whole curve.  And those factors can be estimated and included in regressions, effectively enabling regression on a curve.

In a stimulating new paper, “The Time-Varying Effects of Conventional and Unconventional Monetary Policy: Results from a New Identification Procedure”, Atsushi Inoue and Barbara Rossi put that insight to very good use. They use DNS yield curve factors to explore the effects of monetary policy during the Great Recession.  That monetary policy is often dubbed "unconventional" insofar as it involved the entire yield curve, not just a very short "policy rate".

I recently saw Atsushi present it at NBER-NSF and Barbara present it at Penn's econometrics seminar.  It was posted today, here.

Wednesday, January 20, 2016

Time-Varying Dynamic Factor Loadings

Check out Mikkelsen et al. (2015).  I've always wanted to try high-dimensional dynamic factor models (DFM's) with time-varying loadings as an approach to network connectedness measurement (e.g., increasing connectedness would correspond to increasing factor loadings...).  The problem for me was how to do time-varying parameter DFM's in (ultra) high dimensions.  Enter Mikkelsen et al.  I also like that it's MLE -- I'm still an MLE fan, per Doz, Giannone and Reichlin.  It might be cool and appropriate to endow the time-varying factor loadings with factor structure themselves, which might be a straightforward extension (application?) of Sevanovic (2015).  (Stevanovic paper here; supplementary material here.)

Maximum Likelihood Estimation of Time-Varying Loadings in High-Dimensional Factor Models


Jakob Guldbæk Mikkelsen (Aarhus University and CREATES) ; Eric Hillebrand (Aarhus University and CREATES) ; Giovanni Urga (Cass Business School)


2015


In this paper, we develop a maximum likelihood estimator of time-varying loadings in high-dimensional factor models. We specify the loadings to evolve as stationary vector autoregressions (VAR) and show that consistent estimates of the loadings parameters can be obtained by a two-step maximum likelihood estimation procedure. In the first step, principal components are extracted from the data to form factor estimates. In the second step, the parameters of the loadings VARs are estimated as a set of univariate regression models with time-varying coefficients. We document the finite-sample properties of the maximum likelihood estimator through an extensive simulation study and illustrate the empirical relevance of the time-varying loadings structure using a large quarterly dataset for the US economy.

Thursday, September 24, 2015

Coolest Paper at 2015 Jackson Hole

The Faust-Leeper paper is wild and wonderful.  The friend who emailed it said, "Be prepared, it’s very different but a great picture of real-time forecasting..." He got it right.

Actually his full email was, "Be prepared, it’s very different but a great picture of real-time forecasting, and they quote Zarnowitz." (He and I always liked and admired Victor Zarnowitz. But that's another post.)


The paper shines its light all over the place, and different people will read it differently. I did some spot checks with colleagues. My interpretation below resonated with some, while others wondered if we had read the same paper. Perhaps, as with Keynes, we'll never know exactly what Faust-Leeper really, really, really meant.


I read Faust-Leeper as speaking to f
actor analysis in macroeconomics and finance, arguing that dimensionality reduction via factor structure, at least as typically implemented and interpreted, is of limited value to policymakers, although the paper never uses wording like "dimensionality reduction" or "factor structure".

If 
Faust-Leeper are doubting factor structure itself, then I think they're way off base. It's no accident that factor structure is at the center of both modern empirical/theoretical macro and modern empirical/theoretical finance. It's really there and it really works.

Alternatively, if they're implicitly saying something like this, then I'm interested:


Small-scale factor models involving just a few variables and a single common factor (or even two factors like "real activity" and "inflation") are likely missing important things, and are therefore incomplete guides for policy analysis


Or, closely related and more constructively: 


We should cast a wide net in terms of the universe of observables from which we extract common factors, and the number of factors that we extract. Moreover we should examine and interpret not only common factors, but also allegedly "idiosyncratic" factors, which may actually be contemporaneously correlated, time dependent, or even trending, due to mis-specification.


Enough.  Read it for yourself.


[General note: My use of terms like "factor modeling" throughout this post should be broadly interpreted to include not only explicit reduced-form statistical/econometric dynamic factor modeling, but also structural DSGE modeling.]