Econometrics, economics, finance, random rants.

Econometrics, economics, finance, random rants...
Showing posts with label e-things. Show all posts
Showing posts with label e-things. Show all posts

Sunday, February 5, 2017

Data for the People

Data for the People, by Andreas Weigend, is coming out this week, or maybe it came out last week. Andreas is a leading technologist (at least that's the most accurate one-word description I can think of), and I have valued his insights ever since we were colleagues at NYU almost twenty years ago. Since then he's moved on to many other things; see http://www.weigend.com

Andreas challenges prevailing views about data creation and "data privacy". Rather than perpetuating a romanticized view of data privacy, he argues that we need increased data transparency, combined with increased data literacy, so that people can take command of their own data. Drawing on his work with numerous firms, he proposes six "data rights":

-- The right to access data
-- The right to amend data
-- The right to blur data
-- The right to port data
-- The right to inspect data refineries
-- The right to experiment with data refineries

Check out Data for the People at http://ourdata.com.


[Acknowledgment: Parts of this post were adapted from the book's web site.]

Thursday, December 10, 2015

New Elsevier: Good or Bad?


This just in from Elsevier.  Hardly my favorite firm, but still.  Does it resonate with you, for your own future publications? I am intrigued, insofar as it may actually have scientific value in disseminating research and helping people world-wide to see "seminars" that they wouldn't otherwise see. On the other hand, it would be more work, and the Elsevier implementation may be poor.  (Click below on  "View a sample presentation".  Can you find it?  I looked for five minutes and couldn't.  Maybe it's just me.)  Thoughts?


Elsevier
AudioSlides - explain your paper in your own words

Congratulations on the acceptance of your article Improving GDP Measurement: A Measurement-Error Perspective for publication in Journal of Econometrics. Now that your article is set to be published online, it is time to think about ways to promote your work and get your message across to the research community.
How would you like to present your research to a large audience, highlighting your main findings and articulating the relevance of your results in your own words? With AudioSlides, a new and free service by Elsevier, you can do exactly that!

The AudioSlides Authoring Environment* enables you to create an interactive presentation from your slides and add voice-over audio recordings, using only a web browser and a computer with a microphone. When it's ready, your presentation will be made available next to your published article on ScienceDirect. Click here to get started.

Benefits for authors and readers:
  • Promote your work and summarize your research in your own words
  • Support readers to quickly determine the relevance of your paper
  • Use a dedicated, easy-to-use website to create your AudioSlides presentation
  • AudioSlides presentations can be embedded in other websites
  • AudioSlides presentations will be made available next to your published article on ScienceDirect (View a sample presentation)
  • Did we mention it's free?
We hope you share our enthusiasm about this new service. Thank you for your interest!
Yours sincerely,
Hylke Koers
Head of Content Innovation, STM Journals, Elsevier



*Note that AudioSlides presentations are limited to 5 minutes maximum and should be in English. We advise you to keep the slides limited in number and simple so that they can also be viewed at low resolution (the default width on ScienceDirect for the viewer application is 270 pixels).

AudioSlides presentation are not peer-reviewed, and will be made available with your published article without delay after you have finalized your presentation and completed the online copyright transfer form. See also the Terms & Conditions.

Further information, instructions and an FAQ are available
at http://www.elsevier.com/audioslides.

For questions regarding Audioslides please visit http://help.elsevier.com/app/answers/list/p/8828/c/9413.
Elsevier

Monday, December 1, 2014

Quantum Computing and Annealing



My head is spinning. The quantum computing thing is really happening. Or not. Or most likely it's happening in small but significant part and continuing to advance slowly but surely. The Slate piece from last May still seems about right (but read on). 

Optimization by simulated annealing cum quantum computing is amazing. It turns out that the large and important class of problems that map into global optimization by simulated annealing is marvelously well-suited to quantum computing, so much so that the D-Wave machine is explicitly and almost exclusively designed for solving "quantum annealing" problems. We're used to doing simulated annealing on deterministic "classical" computers, where the simulation is done in software, and it's fake (done with deterministic pseudo-random deviates). In quantum annealing the randomization is in the hardware, and it's real

From the D-Wave site:
Quantum computing uses an entirely different approach than classical computing. A useful analogy is to think of a landscape with mountains and valleys. Solving optimization problems can be thought of as trying to find the lowest point on this landscape. Every possible solution is mapped to coordinates on the landscape, and the altitude of the landscape is the “energy’” or “cost” of the solution at that point. The aim is to find the lowest point on the map and read the coordinates, as this gives the lowest energy, or optimal solution to the problem. Classical computers running classical algorithms can only "walk over this landscape". Quantum computers can tunnel through the landscape making it faster to find the lowest point.
Remember the old days of "math co-processors"? Soon you may have a "quantum co-processor" for those really tough optimization problems! And you thought you were cool if you had a GPU or two.

Except that your quantum co-processor may not work. Or it may not work well. Or at any rate today's version (the D-Wave machine; never mind that it occupies a large room) may not work, or work well. And it's annoyingly hard to tell. In any event, even if it works, the workings are subtle and still poorly understood -- the D-Wave tunneling description above is not only simplistic, but also potentially incorrect.

Here's the latest, an abstract of a lecture to be given at Penn on 4 December 2014 by one of the world's leading quantum computing researchers, Umesh Vazirani of UC Berkeley, titled "How 'Quantum' is the D-Wave Machine?":
A special purpose "quantum computer" manufactured by the Canadian company D-Wave has led to intense excitement in the mainstream media (including a Time magazine cover dubbing it "the infinity machine") and the computer industry, and a lively debate in the academic community. Scientifically it leads to the interesting question of whether it is possible to obtain quantum effects on a large scale with qubits that are not individually well protected from decoherence.

We propose a simple and natural classical model for the D-Wave  machine - replacing their superconducting qubits with classical magnets, coupled with nearest neighbor interactions whose strength is taken from D-Wave's specifications. The behavior of this classical model agrees remarkably well with posted experimental data about the input-output behavior of the D-Wave machine.

Further investigation of our classical model shows that despite its simplicity, it exhibits novel algorithmic properties. Its behavior is fundamentally different from that of its close cousin, classical heuristic simulated annealing. In particular, the major motivation behind the D-Wave machine was the hope that it would tunnel through local minima in the energy landscape, minima that simulated annealing got stuck in. The reproduction of D-Wave's behavior by our classical model demonstrates that tunneling on a large scale may be a more subtle phenomenon than was previously understood...
Wow. I'm there.

All this raises the issue of how to test untrusted quantum devices, which brings us to the very latest, Vizrani's second lecture on 5 December, "Science in the Limit of Exponential Complexity." Here's the abstract:
One of the remarkable discoveries of the last quarter century is that quantum many-body systems violate the extended Church-Turing thesis and exhibit exponential complexity -- describing the state of such a system of even a few hundred particles would require a classical memory larger than the size of the Universe. This means that a test of quantum mechanics in the limit of high complexity would require scientists to experimentally study systems that are inherently exponentially more powerful than human thought itself! 
A little reflection shows that the standard scientific method of "predict and verify" is no longer viable in this regime, since a calculation of the theory's prediction is rendered computationally infeasible. Does this mean that it is impossible to do science in this regime? A remarkable connection with the theory of interactive proof systems (the crown jewel of computational complexity theory) suggests a potential way around this impasse: interactive experiments. Rather than carry out a single experiment, the experimentalist performs a sequence of experiments, and rather than predicting the outcome of each experiment, the experimentalist checks a posteriori that the outcomes of the experiments are consistent with each other and the theory to be tested. Whether this approach will formally work is intimately related to the power of a certain kind of interactive proof system; a question that is currently wide open. Two natural variants of this question have been recently answered in the affirmative, and have resulted in a breakthrough in the closely related area of testing untrusted quantum devices. 
Wow. Now my head is really spinning.  I'm there too, for sure.

Wednesday, June 18, 2014

Windows File Copy

Estimation
Of course we've all wondered for decades, but during the usual summertime cleanup I recently had to copy massive numbers of files, so it's on my mind. Seriously, what is going on with the Windows file copy "remaining time" estimate? Could an average twelve-year-old not code a better algorithm? (Comic from XKCD.)








Wednesday, January 8, 2014

Elements of Statistical Learning: A Stunningly Good Job of LaTeX to pdf to Web


 Click Here  

A very Happy New Year to all! Here's a little thing to start us off.

I happened to be thinking about principal-component regression vs. ridge regression yesterday, so as usual I consulted the Hastie-Tibshirani-Friedman (HTF) classic, Elements of Statistical Learning. Where did I get that gorgeous book pdf? (Look through it; the form is as wonderful as the substance, and see also the similarly-wonderful new James, Witten, Hastie and Tibshirani (JWHT), Introduction to Statistical Learning, with Applications in R.) Both are freely (and legally!) available as pdf on the web. Interestingly, both are also for sale by Springer in the usual ways.

So what's up? In path-breaking arrangements, HTF and JWHT negotiated deals in which they're free to post the book and Springer is free to sell it. And by all accounts the outcomes have been superb for all. Thanks, HTF and JWHT, for promoting best-practice science, and thanks Springer, for doing the right thing.  May many more follow suit.


Monday, December 2, 2013

The e-Writing Jungle Part 3: Web-Based e-books Using Python / Sphinx

In the previous Parts 1 and 2, I essentially dealt with two extremes: (1) LaTeX to pdf to web, and (2) raw HTML (however arrived at) with math rendered by MathJax. Now let's look at something of a middle ground: the Python package, Sphinx, for producing e-books.

Part 3: Python / Sphinx

Parts 1 and 2 of Quantitative Economics, by Stachurski and Sargent, are great routes into Python for economists. There's lots of good comparative discussion of Python vs. Matlab or Julia, the benefits of public-domain, open-source code, etc. And it's always up to the minute, because it's an on-line e-book! Just check it out.

Of course we're interested here in e-books, not Python per se. It turns out, however, that Stachurski and Sargent is also a cutting-edge example of a beautiful e-book. It's effectively written in Python using Sphinx, which is a Python package that started as a vehicle for writing software manuals. But a manual is just a book, and one can fill a book with whatever one wants.

Sphinx is instantly downloadable, beautifully documented (the documentation is written in Sphinx, of course!), open source, and public domain (licensed under BSD). ReStructuredText is the powerful markup language. (You can learn all you need in ten minutes, since math is the only complicated thing, and math stays in LaTeX, rendered either by JavaScript via MathJax or as png images, your choice.) In addition to publishing to HTML, you can publish to LaTeX or pdf.

Want to see how Sphinx performs with math even more dense than Stachurcski and Sargent's? Just check, for example, the Sphinx book Theoretical Physics Reference.  Want to  see how it performs with graphics even more slick than Stachurcski and Sargent's? Just check the Matplotlib Documentation. It's all done in Sphinx.

Sphinx is a totally class act. In my humble opinion, nothing else in its genre comes close.

Monday, November 18, 2013

The e-Writing Jungle Part 2: The MathML Impasse and the MathJax Solution

Back to LaTeX and MathJax and MathML and Python and Sphinx and IPython and R and Knitter and Firefox and Chrome and ...

In Part 1, I praised e-books done as LaTeX to pdf to the web, perhaps surprisingly. Now let's go the other way, to an e-book done natively on the web as HTML. Each approach is worth considering, depending on the application, as each has different costs and benefits.

Part 2: The MathML Impasse and the MathJax Solution

All we want is an HTML version with native support and beautiful rendering of mathematics. That's what HTML5 does, except for a small detail: many browsers (IE, Chrome, ...) won't display HTML5. The real problem is MathML, which is embedded in HTML5, and which is the key to math fonts in HTML5 or anywhere else. It's not just a question of browser suppliers finally waking up and flipping on the MathML switch; rather, successful MathML integration turns out to be really hard (seriously, although I don't really know why), and there are also security issues (again seriously, and again I don't really know why). For those reasons, the good folks at Microsoft and Google, for example, have now basically decided that they'll never support MathML. There's a lot of noise about all this swirling around right now -- some of it quite bitter -- but a single recent informative and entertaining piece will catapult you to the cutting edge, "Google Subtracts MathML from Chrome, and Anger Multiplies," by Steven Shankland.

The bottom line: Math has now been officially sentenced to an eternity of second-class web citizenship, in the sense that native and broad math browser support is not going to happen. But that brings us to MathJax, a JavaScript app that works with HTML. You simply type in LaTeX and MathJax finds any math expressions and renders them beautifully. (For an example see my recent post On the Wastefulness of (Pseudo-) Out-of-Sample Predictive Model Comparisons, which was done in LaTeX and rendered using MathJax.) Note well that MathJax is not just pasting graphics images; hence its output scales nicely and works well on mobile devices too. For all you need to know, check out "MathML Forges On," by Peter Krautzberger.

So what's the big problem? Doesn't HTML plus MathJax basically equal HTML5, with the major additional benefit that it actually works? Of course it's somewhat insulting to us math folk, and certainly it's aesthetically unappealing, to have to overlay something on HTML just to get it to display math. (I'm reminded of the old days of PC hardware, with separate "math co-processors.") And there are other issues. For example, MathJax loads from the cloud (unless it's on your machine(s), which requires installations and updates, and which can't be done for mobile devices), and the MathJax math rendering may take a few seconds or more, depending on the speed of your connection and the complexity/length of your math.

But are any of the above "problems" truly serious? I don't think so. On the contrary, MathJax strikes me as a versatile and long-overdue solution for web-based math. And its future looks very bright, with official supporters now ranging from the American Mathematical Society to Springer to Matlab. (Not that I'm a fan of Matlab any longer -- please join the resistance, purge Matlab from your life, and replace it with Python and R -- but that's a topic for another day.)

[Next: Python, Sphinx, ...]

Thursday, November 7, 2013

The e-Writing Jungle Part 1: LaTeX to pdf to the Web

LaTeX and MathML and MathJax and Python and Sphinx and IPython and R and Knitter and Firefox and Chrome and ...

My head is spinning with all this stuff. Maybe yours is too.

One thing is clear: The traditional academic book publishing paradigm (broadly defined) is cracking and will soon be crumbling. In the emerging e-paradigm there will be essentially no difference among books, courses, e-books, e-courses, web sites, blogs, and so on. With no loss of generality, then, let's just call it all "e-books," filled with text, color graphics, audio/video, animations, interactive learning tools, massive numbers of internal and external hyper-links, etc.

An interesting question is how to create (``write"?) and distribute such e-books. The amazing thing is that the answer remains unclear. Both pitfalls and opportunities abound. Here are some thoughts.

Part 1:  LaTeX to pdf to the Web

One obvious e-book creation and distribution route is traditional LaTeX, compiled to pdf and posted on the web. Effete insiders now sneer at that, viewing it as little more than posting page photos of an old-fashioned B&W paper book. I beg to differ. What's true is that most people still fail to use the e-capabilities of LaTeX, so of course their pdf product is little more than an e-copy of an old paper book, but that's their fault. All of the above-mentioned e-desiderata are readily available in LaTeX/pdf/web; one just has to use them!

Moreover, LaTeX/pdf/web has at least two extra benefits relative to a website (say). First, trivially, the pdf is instantly printable on-demand as a beautiful traditional book, which is sometimes useful. Second, and more importantly, the linear beginning-to-end layout of a "book" -- in contrast to the non-linear jumble of links that is that is a website -- is pedagogically invaluable when done well. That is, good authors put things in a precise order for a reason, and readers benefit by reading in that order.

OK, you say, but how to restrict access only to those who pay for a LaTeX/pdf/web e-book? (It's true, a pdf web post is basically impossible to copy-protect.) My present view is very simple: Just get over it and forget the chump change. Scholarly monographs and texts are labors of love; the real compensation is satisfaction from helping to advance and spread knowledge. And if that's not quite enough, rest assured that if you write a great book you'll reap handsome monetary rewards in subtle but nevertheless very real ways, even if you post it gratis.

[To be continued. Next: HTML and MathML and LaTeXtoHTML5 and MathJax and ...]

Friday, November 1, 2013

LaTeX/MathJax Rendering in Blog Posts

It seems that LaTeX/MathJax is working fine with my blog, including with mobile devices, which is great (see, for example, my recent post On the Wastefulness of (Pseudo-) Out-of-Sample Predictive Model Comparisons). However, a problem exists for those with email delivery, who just get raw LaTeX dumped into the email. If that happens to you, simply click on the blog post title in the email. Then you'll be taken to the actual blog, and it should render well.

Please let me know of any other problems!

Monday, October 7, 2013

Why You Should Join Twitter

Sounds silly, but it's not. I got talked into joining a few weeks ago, and I'm glad I did. I rarely tweet (except to announce new No Hesitations posts), but I follow others. Several times in the last few weeks alone, various pieces of valuable information arrived. Great stuff.

FYI here are some random things that I'm currently following. (In total I follow about 25, but for some reason Google Blogger crashes if I try to paste them all here.)

[By the way, I will soon stop posting announcements of new blog posts to Facebook groups, instead announcing exclusively with a tweet. SERIOUSLY. So join Twitter and follow @FrancisDiebold.]

Saturday, October 5, 2013

Pure Brilliance From FRB St. Louis: EconomicAcademics.org

This just in from Christian Zimmermann and the RePEc Team at FRB St. Louis:

"Congratulations, you made the list! .. The Federal Reserve Bank of St. Louis is launching a blog aggregator, EconomicAcademics.org, to highlight and promote the discussion of economics research. Your blog is part of this effort. This email explains why and how you can help promote the discussion of economic research in the blogosphere ... EconAcademics.org lives at http://econacademics.org/ and aggregates blog posts that discuss economic research. The aggregator looks through blog posts for a link to some research indexed on a RePEc service, currently EconPapers, IDEAS and NEP. IDEAS then also links back from the abstract page to the blog posts ... This blog aggregator is provided by the Federal Reserve Bank of St. Louis, which also offers with FRED database and graphing tool as a useful resources for bloggers. Feel free to use the graphs on your blog, best done by embedding them so that readers can click on them to get more details about the data. FRED lives at http://research.stlouisfed.org/fred2/." 

This is totally brilliant. First, it's a brilliant public service. I am grateful. Everyone should be grateful. But second, and this is really what I want to emphasize, ya gotta love the brilliant business/marketing move. Instantly, every blogger now has a strong incentive to report on (and link to) RePEc papers whenever possible -- in case you missed it above, blog posts get noticed by the aggregator only if/when they link to a RePEc paper -- and hence authors have a correspondingly strong incentive to put their papers on RePEc. And it's all tangled up with the wonderful FRED. The idea may not make billions for FRBSL/RePEc/FRED, but in its own way it's as brilliant as Google's pagerank (and cynics will say as obvious -- just sour grapes).

Oh wait. I forgot to mention a RePEc paper above, so this post won't get picked up by EconAcademics. Hmmm... In the future I'll have to change that...

Anyway, SSRN et al. must be reeling! Of course it will be interesting to see how they and others respond. FRBSL/RePEc/FRED have scored a significant first-mover blow, but surely the fight isn't over. And as usual with healthy competition, everyone will benefit.

Friday, September 6, 2013

Tom Sargent, Quantitative Economics, and Python

Speaking of Tom Sargent, check out his latest at http://www.quant-econ.net/.  (Thanks to Frank DiTragila for forwarding a few days ago.)  Python features prominently...

Friday, May 31, 2013

Research Computing / Data / Writing Environments

You asked for favorite blogs, and I obliged. You also asked for favorite computing environments etc., so here goes. I look forward to comments telling me why I'm wrong, stupid, insane, or worse.

At some level, who cares about computing etc.? If you're deep into retirement, a SAS guru writing in WordPerfect (say), is it worth updating? Almost surely not.

But if your investment horizon is longer, and if you want to be on the cutting edge, and if you want my opinion, I certainly have one. What follows is in part prescriptive, although I realize that one size surely can't fit all. In any event it's certainly descriptive; it's basically what I do, in principle if not always in practice. (I admit that I'm still rather fond of certain ancient low-level environments like Fortran, and certain high-level environments like Eviews.)

My computing epiphany of recent years centers on R, a mid-level environment.  I find that R is most often the place to be.  Check R Studio IDE, a wonderful R work environment. Check R-bloggers (I should have mentioned it as a favorite blog -- thanks to Frank DiTraglia for reminding me). And for all you parallelization freaks with GPU's, check the CRAN Task View on High-Performance and Parallel Computing with R and the R Tutorial.

For time-series data, check Quandl. Totally amazing. Just click on the link and see for yourself. (And yes, there's a seamless R interface.) Imagine having basically any time-series you could ever want, instantly available and continuously updated, for use in your R code.

For writing, obviously it's LaTeX. My favorite flavor is MiKTeX. Enough said.

Now here's the first kicker. I already mentioned that Quandl and R are interfaced. But so too are R and LaTex, via Sweave. So now data, computing and writing are all linked. Imagine writing a book (in LaTeX) whose graphics and statistical analyses (in R) are automatically updated in real time as new data arrive (in Quandl). It's not a dream.

And here's the second kicker. Everything I've emphasized is public domain, open source, free. Who says that you get only what you pay for? This is highest quality everything, cutting edge, with no license hassles, no renewal hassles, no payment hassles.

Power to the people!