Computable Document Format
wolfram.com
wolfram.com
So, with some knowledge of the man, I can say honestly that the charge that he's oblivious of the evolution of HTML is completely laughable. A random illustration: Wolfram Research was one of the first companies to go online in the early 90s (as it happens, Tim Berners-Lee is a long-time Mathematica user). An amusing story: it was also one of the only companies to survive the original Morris worm unscathed, owing to deliberate use of obscure Japanese computers for WRI's gateway.
While I can't talk about unannounced technologies, I can say that HTML5 plays a pretty crucial role in our future technologies. In fact, CDF will eventually have a server-side incarnation that relies on HTML5 for client-side interactivity.
If someone posts a link in their profile and you don't follow it, they can't be held accountable.
I sincerely hope he's also dead wrong on the acceptability of a single company owning your communication in 2011.
I don't think you looked very closely at CDF. It looks like it has a very large fraction of Mathematica in the player. You aren't going to get anywhere close to what you can do in CDF via HTML5 and JavaScript without a truly ungodly amount of work, and it would still be unusable for many things due to performance issues.
The same argument was being made in favor of Flash a few years ago. Eventually it turns out that someone does that 'ungodly amount of work.' Why shouldn't it be Wolfram? He would have a leg up on everyone if he embraced the new standards and opened up the Mathematica walled garden a little.
I don't think so. I think it's more likely that he's fully aware of said evolution but appreciates that most people are not developers and want a nice self-contained live document format that they can email as a single-file attachment if need be.
Most people don't have a copy of Mathematica lying around either, and it appears to be a requirement for authoring CDFs: http://www.wolfram.com/cdf/faq/ .
Look, I write and create multimedia content, I like powerful composition and editing tools. But managing all the structural information and assets for a large compound document or media project is a lot of work, the kind of work for which I prefer to be paid or rewarded in kind. When I'm just consuming and sharing the work of others, then I don't want to do all that work and I prefer a nice self-contained package that doesn't impose any administrative overhead. When it's as easy to store or share a HTML+CSS+JS document, online or off, and have it appear exactly the same way in a completely device-independent manner or even as the output of a printer, then I'll sing its praises. What I do not want is a bunch of extra files to keep track of for a single document that I just wish to add to my library and might not open again for a year. I am often much more interested in the content of a document than in the ability to edit, deconstruct, or radically reformat it.
When I was younger, I cared more about having control over things like typesetting, page flow, and other design issues, and also cared more about everything being as editable as possible. Now that I'm older I'm far more concerned with what a document is about than with how it looks. So I tend to open Picasa ten times for every time I launch Photoshop or my Camera RAW editing software, and I tend to read PDFs in the browser or in Acrobat reader far more often than I run up the full Acrobat environment to make a PDF file.
I suggest that you focus less on how you would do things differently if you were Stephen Wolfram, and more on whether his CDF format might open up some new economic opportunities, the way that PDF files have leveraged simplicity into ubiquity.
I'm suggesting that Wolfram is trying to reproduce the success of the PDF and Flash business model at the very moment that that model is in decline.
Similarly, Flash has hardly gone away. Of course, 1) Flash has had persistent stability and performance problems, 2) Apple is waging a very public battle against it via iOS devices, and 3) its only real killer app, streaming video, is being folded into several (competing) standards. But it still has massive penetration.
The thing is, Mathematica is too big and too rich to ever achieve standardization. If you want the ability to easily inject graph theory, computer vision, symbolic statistics, non-linear optimization, discrete math, control theory, etc etc into your documents, CDF is unmatchable.
Now, without seeing with your own eyes the kinds of crazy stuff you can do with Mathematica, you might well be skeptical, but the next few months will change your mind, I promise.
You need to disclose the fact you work for the company that makes Mathematica.
I also disagree with this whole notion: There's no upper limit on what can be standardized. Unless you're implying a patent portfolio or some other anti-competitive practice, of course, which is another notion entirely.
It seems that all we need is an extension, I propose ".htz" and a mime type, say "application/x-html-bundle". Browser's would just download it and open the containing "index.html" (at the top level or inside a unique top level directory) and open other assets relative to that.
As far as I understand the story, Stephen told Matthew: here, this CA seems rich enough for universality, can you prove this for the book? Matthew duly proved it -- a tour-de-force proof, to be sure. Then Matthew broke his NDA by publishing the result early. Stephen litigated to avoid it becoming public prematurely (it was, as you say, the center-piece of the book).
I'm not privy to all the ins and outs and what-have-yous, but that seems fair. If you tried to publish something behind your advisor's back in an academic setting, you'd have a lot to answer for.
Matthew eventually did publish the Rule 110 proof under his own name in Wolfram's own journal.
I speak for myself, not WRI, here.
If you tried to publish something behind your advisor's
back in an academic setting, you'd have a lot to answer for.
Hmm, but similarly if I made a contribution to a breakthrough, and my adviser was going to publish it without my name being associated with it at all, I would feel cheated. I think that's probably a better analogy.It is a very strange setup where Stephen is claiming he created the work done by others. Legal, certainly, but not very nice. I guess I wouldn't care to work for one of the most egotistical people on earth, however smart he is.
"His initial results were encouraging, but after a few months he became increasingly convinced that rule 110 would never in fact be proved universal. I insisted, however, that he keep on trying, and over the next several years he developed a systematic computer-aided design system for working with structures in rule 110. Using this he was then in 1994 successfully able to find the main elements of the proof. Many details were filled in over the next year, some mistakes were corrected in 1998, and the specific version in the note below was constructed in 2001."
I had read this well known review:
http://cscs.umich.edu/~crshalizi/reviews/wolfram/
"The real problem with this result, however, is that it is not Wolfram's. He didn't invent cyclic tag systems, and he didn't come up with the incredibly intricate construction needed to implement them in Rule 110. This was done rather by one Matthew Cook, while working in Wolfram's employ under a contract with some truly remarkable provisions about intellectual property. In short, Wolfram got to control not only when and how the result was made public, but to claim it for himself. In fact, his position was that the existence of the result was a trade secret. Cook, after a messy falling-out with Wolfram, made the result, and the proof, public at a 1998 conference on CAs. (I attended, and was lucky enough to read the paper where Cook goes through the construction, supplying the details missing from A New Kind of Science.) Wolfram, for his part, responded by suing or threatening to sue Cook (now a penniless graduate student in neuroscience), the conference organizers, the publishers of the proceedings, etc. (The threat of legal action from Wolfram that I mentioned at the beginning of this review arose because we cited Cook as the person responsible for this result.)"
That doesn't put things in a very good light. I have no idea if it is accurate or not. I just found your initial attempt to paint Cook as the bad guy unfortunate. Why would Cook have tried to publish on his own? That would be very unusual, and the simplest explanation is that it must have been some sort of disagreement over academic credit. Why else would he do that?
At least it's an open format apparently: http://www.wolfram.com/cdf/faq/#aboutcdf
Given that CDF files seem to be able to do most things Mathematica can do, I'm guessing that might mean "publically available to anyone who pays us for a Mathematica SDK licence".
Sure, but where is it?!
Good luck finding it: http://www.google.com/search?q=Computable+Document+Format+sp...
I can see big corp. getting excited over this just like 99% of the crap they buy and don't use.
That said, this doesn't exactly move forward with global progress on the problem or solution (as well it shouldn't wolfram is a corp after all... in the business to make money).
On the other hand, if you are, like me, one of the many people in academia who would love to have a way to embed interactive graphs and tables into your papers to make them more understandable for the reader, then this is potentially very useful. Especially since very few people in academia (outside CS departments anyway) know or have any interest in learning HTML5, and many already use or have easy access to Mathematica to begin with.
Also, Wolfram being a corporation is irrelevant. Some problems are best solved by large corporations trying to make money. That may or may not be the case here, but they definitely have an incentive to keep the format open (in some sense) and promulgate the free CDN reader (much like Adobe has done in the past).
For sure, web developers and programmers can do most of what CDF can do, on their own, in HTML5. It might take them a lot longer, but they could certainly end up with a nice finished product.
This doesn't solve the problem. The problem is neatly illustrated by the fact that news organizations, which have a huge incentive to make compelling, sticky interactivity that wraps their news properties, haven't gone for it. I've only seen two non-trivial uses, the NYT and BBC News, and its clear these were bespoke jobs that cost them a lot of money.
The same goes for textbook publishers, scientists, NGOs, etc, anywhere were technical communication could be significantly improved with interactive documents. This problem remains unsolved.
CDF aims to make it possible for someone to crank out an interactive figure or document in a matter of hours, not weeks, with very little code.
A side comment: I say this without any real proof, but WRI specializes in doing interesting things that are economically self-sustaining, rather than things that make a lot of money. Mathematica is far from a cash cow, and WRI is a small company (~500 people), but it has lasted 25 years, and it regularly adds cutting edge technology to its portfolio. Obviously, it gets to balance profit with "interestingness" mainly because it is privately owned, and Stephen likes collecting interesting people and interesting projects for them to do.
I've thought about this for a while now. It seems terribly silly for us to spend so much time formatting our papers as pdfs when people mostly read them on a computer anyway. I've been to conferences that - smartly - don't even print the proceedings. They just hand you a USB key with all of the papers and a table of contents in HTML that has authors and titles that links to the correct file.
I have a few thousand pdf files on my hard disk and already find it hard to manage the collection. If every single on of them had a source tree it would make me cry.
There's another comment downthread where I talk about why I don't actually want all the possible editing capabilities to be available most of the time. Sometimes you want to maximize convenience rather than control.
One (of the many i'm sure) problems here you defined are: "...who would love to have a way to embed interactive graphs and tables into your papers to make them more understandable for the reader..."
The solution however may not necessarily be a private 'reader', 'player', or 'binary'. Though the 'format' may be 'open' this single implementation isn't.
You bring up Adobe which isn't a very good business case to follow as the very reasons I brought up are pressing the industry to move toward open implementations[1][2], away from 'free' binary distributions and plugins (for a various list of reasons from security[2], to access on other systems/apps).
This really isn't 1999, and CDF isn't yet a popular nor a de-facto standard[1] much like PDF was. Besides you and I are much more capable these days with newer and _open_ technologies are we not (browser, linux, open documents, etc...)?
Speaking of documents, other examples of 'perceived' open standard files are Microsoft's ill faded OSP[3] promise which spawned traction for OD/F[4] and other.
These reasons are why HTML5/CSS3/JS as a basis to create an open two way street for 'documents' and/or formats are so powerful. It is not enough to only provide an 'open' format, but also an open implementation. This way both use of and implementation of such product be beneficial toward academic progress. Why wouldn't that alone be worth it?
Thus a likely more popular solution I am proposing to your problem could be a service that uses HTML5/CSS3/JS in the delivery which solves your problem in a WYSIWG general user manner. Especially so as the very tools (your browser, and a million libraries must I really list them all?), UIX experience, and entrepreneurs (HN! Y-comb!) already exist!
To your quote "Some problems are best solved by large corporations trying to make money" I would say the same to "Some problems are best solved by small groups of entrepreneurs or open source developers trying to make money and/or looking for peer fame."
I would go further in saying that for 'open' standards and implementations, that small group of entrepreneurs or open source developers are the ones carrying the torch of open-ness [5][6][7][8][and on and on...] and innovation.
[1] "was originally a proprietary format controlled by Adobe", http://en.wikipedia.org/wiki/Portable_Document_Format
[2] http://andreasgal.com/2011/06/15/pdf-js/
[3] http://en.wikipedia.org/wiki/OpenDocument
[4] http://en.wikipedia.org/wiki/Microsoft_Open_Specification_Pr...
[5] http://blog.documentfoundation.org/2011/06/01/statement-abou...
[6] http://www.infoworld.com/d/application-development/oracle-ha...
If anyone would like to join up and create an alternative I'm game. Lets do this.
My full info is in my profile.
Having looked at the technology in depth I can appreciate the notion of building into the document the computation that went into it. This could be a killer way to distribute component data sheets where all the graphs were 'live'.
That being said, it scares the crap out of me. Why? Because I've got Frame documents which are unusable (there is no available version of Frame which will read them, and no legal way to obtain said version) This is a particularly insidious form of bit rot. I save PDF documents on CD with a self contained C language implementation of a PDF reader that can read them and a set of fonts that work that reader.
Without the equivalent for CDF I worry about having critical information (or simply relevant information) that cannot be viewed or used. At least with PDF if you print it out the paper version is still usable. Not so if you don't have an open source version of Mathematica to interpret what you read.
I guess I'll be sticking with org mode though and trying Tangle (http://worrydream.com/Tangle/) as suggested elsewhere on this page.
In other words, possibly nothing, but your original question may or may not be the right question to ask.
Is it me, or is this a sign the format isn't going to be open in any meaningful sense?
One of my side projects is a HTML word processor, and every time I look at a common Word processor feature, HTML offers a ton of free and easy to implement ways to make that feature better.
Dynamic @ Graphics @ Point @ ImageKeypoints @ CurrentImage[]
What does this do? It takes a video stream from your webcam, extracts SURF keypoints ("salient" parts of the image, like corners), and plots them live. Note that all those functions are online at reference.wolfram.com.
THAT'S A CDF! You have access to some heavy duty algorithms covering a vast number of fields. In fact its kind of surprising it fits in a few hundred megs. You're essentially getting the whole of Mathematica, for free.
Give me this:
stream = OpenVideoStream(device)
render1 = stream.surf() render2 = stream.sift() etc.
display(outDevice, stream)
If this doesn't take off, Wolfram dumps the project, removes the software from its website and other channels, and the documents are worthless.
If I invest in this, I'm betting Wolfram is going to boil the ocean or is going to be cost-insensitive regarding how much it takes to keep this available. Neither of those sound like good bets.
Yes, I've been saying that since about 1992. I'm not disagreeing with you, mind; I really like the idea of a lightweight WYSIWYG tool that lets me create complex documents without having to fiddle about with code.
There are many marketing lies in it! For one, it claims that PDF is not embedded in the browser and CDF is yes? "Full page within a browser" is yes but not for JS? "Dynamically interactive charts and diagrams" is partially supported??
Could anyone please explain to me what I am missing here?
http://www.wolfram.com/cdf/compare-cdf/how-cdf-compares.html