Algorithm Development is Broken
blog.algorithmia.com
blog.algorithmia.com
Is this really the case? When I have to solve a particularly unusual or tricky problem, I often have a look at relevant academic papers and go from there – either finding existing solutions, or basing one on the contents of the paper. I kind of assumed most developers did the same.
We are not trying to monopolize basic algorithms. In fact in a free market I don't see how that could happen, since anyone else could come along and implement the same basic CS algorithms and offer a lower price.
The thing that makes us different from say Wikipedia or Github is that we don't provide just source code, but live infrastructure delivering the algorithms as a service. So rather than uploading code to github, and putting the burden on a user cloning the algorithm, setting up virtual servers, deployment systems, load balancers, etc. We handle all that. That way algorithm developers are free to focus on algorithm code, and the infrastructure is taken care of. This infrastructure needs to be paid for however, so our general plan is to charge for using algorithms through the API. It will be up to the developers whether to share source code or not, since they own the IP.
This type of system will encourage these folks to release their source code in a form compatible with your API, and 3rd parties who want to leverage their work will need to either cut-and-paste it to incorporate it into their own system (a nightmare), or access it through a slow network API, likely paying for it along the way. Libraries have solved the problem of code re-use of algorithms for the last several decades. You're basically proposing that LAPACK and BLAS and Mahout and all these great algorithmic libraries would have been best served up without any centralized coordination or binary distribution, but with each individual algorithm behind a separate API with each author publishing each one individually. I don't see how if the goal is really to enable wider distribution of good algorithmic work that you can't see that an heterogenous API-centric approach to this problem (while one that would certainly make you a pile of cash if it catches on) wouldn't really be a great step forward for the industry compared to the centralized open source library model we have now. Good algorithmic work is hard, and the model we have now results in battle hardened, robust, agreed-upon implementations of well understood algorithms.
If you really want to make me a believer, all algorithms that your API supports that have public source code should also be integrated and built into a binary library that a user can download instead. At the very least, the path for a contributor to take what they have written and additionally publish it as a appropriate binary package themselves should be minimal. For example, if I implement something in Ruby that uses your service, I should be able to push it as a un-coupled, reusable gem in a straightforward manner.
The first language we implemented was Java/Scala, largely because of the excellent libraries like Mahout and Weka easily available in Maven Central. But even though the code is there in maven, it is still not trivial to deploy those algorithms as a service. We are trying to make that easier, while at the same time supporting the developers of those libraries in the process.
Being able to run the code locally is in the plans. For the most part, the algorithms in our system are already simply maven packages, or ruby gems, npm packages, etc. Being able to distribute them in binary form is definitely possible. We're still figuring things out right now, but our focus is always on doing right by the algorithm development community.
The 9th DIMACS Implementation Challenge from 8 years ago (see http://www.dis.uniroma1.it/challenge9/ ) served as a way to gather the best algorithms for the shortest path problem. I don't see how this project can be significantly better, or even close to, a traditional effort like that.
It doesn't have to be better at producing good algorithms. It just has to be better at making those algorithms implementations easy to use.
But I still don't think the business model is viable.
The Boost Graph Library includes three different shortest path implementations, and I have that already downloaded. That library likely fits my needs.
In the DIMACS challenge papers at http://www.dis.uniroma1.it/challenge9/papers.shtml you'll see "Single-Source Shortest Paths with the Parallel Boost Graph Library".
Indeed, the Parallel Boost Graph Library library is easily available, and contains two parallel shortest-path implementations.
The flip question is, what if it doesn't fit my needs? Well, then the question is "what do I need"?
The other DIMACS challenges include 1) pre-computation, which is useful for road-based graphs, 2) supercomputer parallelism to search billions of nodes, 3) out-of-core graph searches, for when you don't have enough RAM, and 4) k-shortest path search.
It's very unlikely that this proposed algorithm service provider will implement these alternate algorithms.
Maybe we need more crowdfunding of open source and independent researchers that are independent of academia. I really think open source hardware is crying out for a project and a community to revolutionize how a community can participated with the mostly closed EDA / silicon stack, because frankly, I don't trust anything I can't get the source code to.
Virtual servers? deployment systems? load balancers?!? Huh? It all looks like gibberish to me. What does any of this have to do with developing implementations of new algorithmic techniques?
Your message is vague as vague can be. It also doesn't help that the word "algorithm" is sufficiently general to cover any piece of code.
We want smarter machines, and that means making algorithms more widely available. To achieve that, we're creating a community around algorithm development, where people can contribute code and collaborate with others on that code. Github has done a great job of this, but it is not always trivial to get from a github repo to live code in production. We think that creating a built-in market for algorithms-as-a-service could be a way to reward algorithm developers for their work, and make algorithms easily accessible via REST API.
The types of algorithms we are talking about is not limited by us. My opinion is that the most useful algorithms in the near future will be related to: machine learning, classification, analysis, and optimization. There are also numerous applications for things like image recognition, speech-to-text, and translation, to name a few.
Good luck... that seems like a really intimidating area to learn and a collection of ready to try algorithms could be super useful.
Goal: It's hard to see algorithms innovating except for large scale microoptimizations using higher levels of abstractions, because many everyday purposes have already been solved using whatever the developer chose to ship a product.
Value: For-profit research is a really hard model to scale. Well-respected universities already do this and are more of a factory for talent (via publishing) rather than IP, so it becomes a very indirect horse bet for a business to invest in. In a sense, this would be building a double-ended marketplace of investors and researchers. Also really hard to scale.
Well, this site is apparently a marketplace for algorithms.
If anyone is more interested in open Wikis, I found what two sites using Google:
* RosettaCode.org (code in many languages): http://rosettacode.org/wiki/Sorting_algorithms/Merge_sort
* Algorithmist.com: http://www.algorithmist.com/index.php/Main_Page
I can't tell you how many times I've looked up something there over the years (complete with links to sample code!).
It's not immediately clear on the linked page, but the "By Language", "By Problem", "Algorithm Links" targets will drop down with links to a specific page. The "By Problem" targets link to a page that is very similar to what is in "The Hitchhiker's Guide to Algorithms" part in Skiena's "The Algorithm Design Manual" textbook. They each lead with an image representation of the input and output of the general case of the algorithm, a problem description, some short text description, and then links to implementations. It's not as detailed as what is in the textbook, and I'm not sure how up to date it is these days, but it's a good place to browse about.
People really need to be aware of this site. It's fantastic. His book is great as well.
* Sorting Algorithm Animations: http://www.sorting-algorithms.com/
* Data Structure Visualizations (interactive): http://www.cs.usfca.edu/~galles/visualization/Algorithms.htm...
* Sorting algorithm visualization (interactive): http://sortvis.org/
* 15 Sorting Algorithms in 6 Minutes (video, turn off sound): https://www.youtube.com/watch?v=kPRA0W1kECg
* Algorithm Visualization Portal: http://algoviz.org/avcatalog
* and I just found an old HN discussions with more resource links: https://news.ycombinator.com/item?id=1511332
We strongly believe in open source, and the source code will generally be available (although we may leave that up to the individual algorithm developers to decide). We are algorithm developers ourselves and our goal is to see more intelligent algorithms used as widely as possible.
How is data for the algorithms provided? A lot of times it is big. And messy, and proprietary. There seems to be an implicit assumption that you can just plug different algorithms to different data sets. But I can't think of a project where that has been the case.
I also doubt that "algorithms" are the bottleneck in a lot of projects. I'm not an expert, but I have some personal experience to back up what people say about "data trumping algorithms" (e.g. Peter Norvig and others have written about this)
I would like to hear about some more concrete examples / success stories. "Algorithms" is just too vague. I think if this becomes successful, it will be by first narrowing it to a particular domain, and then generalizing it again.
The data question is huge for us. We have built our system on top of Hadoop + Spark in order to handle large amounts of data, and be able to apply computations to it. You're absolutely right that data is very heavy, and often the limiting factor, so we are doing everything we can to both get data into the system (currently with a focus on streaming data, since that gets around the heavy data problem), and getting algorithms to where the data is.
As far as specific algorithms, there are many that could apply. My personal experience is in machine learning algorithms, so I'm personally biased toward those. There are numerous algorithms like Classification, Clustering, Optimization, Anomaly Detection, etc. which are very CPU-intensive. We will be doing a blog post soon with a demo of live algorithms already in our system to give more concrete examples.
For example, a couple of years ago I worked on an algorithm to find the maximum common subgraph of a set of 2 or more molecular graphs. More specifically, I wanted the largest subgraph in M of N graphs (M<=N), I wanted to define the atom match criteria, and I wanted to require that rings not be broken in the subgraph. (Chemists love rings.)
I did it the old-fashioned way. I read papers, I investigated similar systems, I implemented various implementation details, and I did a lot of testing.
How would I find such an algorithm using this system?
There are only a few people who develop this sort of algorithm. Why might I expect that this system is a better resource than traditional means?
Lots of domains share the same mathematical abstractions, but we don't know because (almost) nobody is trying to discover them. We can know that it's common, because once in a while people do look, and always come back with new shared abstractions.
If somebody created an index that took different notations into account, that'd revolutionize research. But I don't think this system does that.
Of course there are many different notations. One need only look to Newton and Leibniz notations for calculus, or Schrodinger's wave equations and Heisenberg's matrices for quantum mechanics, to know there's a long history of different mathematical abstractions for the same concept.
In my own field, I recently discovered (by reading the literature and getting feedback from conferences) that there appears to be a previously unknown connection between this maximum common subgraph problem and frequent subgraph mining, so it's not like connections are altogether rare.
Indeed, I'll argue that these occur all the time, and research libraries exist in part in order to help people find them.
Which is why I asked why this project might be better than traditional means, and I gave an real-world example to give a basis for discussion.
[1] There's plenty of MCS research, but since the problem itself is known to be NP-hard, most of the mathematical work shows that certain limit cases, like planar graphs or outerplanar graphs, is solvable in polynomial time. As the structures I deal with don't all fall into those categories, I can't use them as a general purpose solution. [2]
[2] I might be able to use them as a special purpose solution, and this project might help identify codes I can use for this case. I just can't figure out a way to make it easier than the current methods.
Why would anyone write about not finding a shared abstraction? You have a significant publication bias there.
The thing with algorithms is that you really have to think. 200 lines of code might take two months to really grok. Especially because the people in the field make certain assumptions when they start, and sometimes even just one of those assumptions takes two weeks to understand and research.
You can't just jump into PhD level research and expect to understand it right away.
I don't know if algorithmia will help solve this problem, but I wish them all the best of luck. Getting actual code next to research is super important and useful.
People tend to focus on initial algorithm development, because that is the academically prestigious intellectually stimulating bit, but the job is really about how to turn algorithms into cash; a much broader problem than the narrow slice that people typically obsess over.
A huge (and frequently overlooked) part of that job is the communication and coordination role between business development and algorithm development. The volume of communication and level of detail required cannot be overstated.
Another huge part of the job is actually turning a piece of research into a functioning product. Whilst a large part of the OP's proposition is intended to address this problem, (kudos) I think that his solution falls short in a big way: It omits the largest part of the solution, which is where the business learns about the algorithm and how the behaviour and performance characteristics of the algorithm interact with the business' problem domain. I.e. how does the business build sufficient expertise and knowledge of their product to be able to effectively sell it. All of these are human problems, mainly oriented around communications and learning.
Having said all that, I would like to encourage the OP in his efforts. I think that it is something that is worth doing, and I really hope he is able to build a business around this idea. I think that technology can help support all of these activities, and this is actually something that I have wanted to do myself for very a long time, so Kudos to the OP for actually taking the chance, going out and doing it!
However, if you want to do something non-trivial and you want it done right, hire a computer scientist or mathematician. No amount of crowd will help if no one in the crowd has a clue what they're doing.
2) Same question as above, but for processing times. Some algos are aimed at operating on the fly and might take ~1 second or less. Others (many that I deal with quite often), might run for days, or even weeks. What sort of processes are you optimizing for, and are long running procs on the radar?
It is an interesting idea, but there is a high bar when competing against my language's package manager, and the 10s of thousands of "algorithms" already out there. Best of luck!
That said, I think there are many algorithms that will work as standalone algorithms. Lots of machine learning type algorithms are mostly CPU-bound. There are companies like http://www.kooaba.com that do image recognition as a service over HTTP. Also things like Siri and Google Now seem to work well enough, despite network latency. Thanks for the feedback!
A non-academic community based around discovery, discussion and implementation of algorithms would be an interesting place to hang around. A MVP would be, basically, a discussion forum with an index of github repos for the implementations, but there are other values to be added on top.
Good luck. I'll be paying attention to what becomes of your project.
1. There are 14 libraries for technical algorithms. 2. We need a meta-library so that we don't need to worry about having so many libraries! 3. There are 15 libraries for technical algorithms.
For example one day I was interested in implementing HyperLogLog(a set cardinality measure that is useful in data analysis). In about 10m I had all the relevant papers on hand and after skimming them I had a pretty good sense of how to implement it.
Similarly if I want to know how to implement a program dependency graph for doing program analysis I can go read a few pages of a paper and get a good description on the algorithm I would need to construct such a thing. I can believe the argument that some of these things are poorly indexed but even a bare minimum of Google searching usually results in useful algorithms. I would argue that often times many of these research algorithms have a bunch of different design decisions that are best explored in the academic literature around them, and an implementation and a few notes is not sufficient exploration.
For example I was recently implementing Paxos and there were tons of little details to be extracted from the papers around that had a big impact on the actual implementation we ended up with. The 'Paxos Made Live' paper from Google had many details that were only relevant/true because of engineering decisions made by the team. If one was presented you with an implementation derived solely from that paper there are multiple incorrect assumptions you could derive.
An instance of this is made apparent in Paxos Made Live. Google essentially fixes their proposer because the have used Google specific details about the number of participants and their availability. The result is that they direct all traffic to a single node, and don't spend a much time talking about leader/proposer selection (which could be useful to your needs).
I also don't buy that an important part is getting the algorithms running as a service. Most so called "algorithms" are nothing more than a subroutine that is need as a piece of a greater whole. I would venture most useful "algorithms" are most likely container data structures and algorithms that operate over them. It seems that these are probably most useful to have as a library. Many libraries have already taken this approach LLVM (algorithms for code generation, albeit not always everything you want), OpenCV for computer vision routines, BLAS for linear algebra, NLOpt for non-linear optimization, and I'm sure one could come up with many more examples of democratized algorithms.
1. Make open source releasing of algorithms compulsory, not optional.
Think of it from your prospective clients' perspective: the upside of digging out an algorithm from academic journals and implementing it yourself is that you get to fully understand the source code (since you end up writing it yourself); you can attest its correctness; and you get to stand on the shoulders of giants by tweaking and extending it later, if you wish.
Admittedly this might not be seen as that much a benefit for business users, but for many actual users of advanced algorithms in the scientific computing community, having to use proprietary algorithms with restrictions on their use has been seen as a significant step backwards, as manifest during the controversy brewed over the "Numerical Recipes" controversy [1,2,3]. Even if the benefits are more a matter of principle than mere practicality, the palpable distaste for proprietary algorithms in the scientific computing community is something you should at least keep in mind, lest you risk alienating a core user base for your product.
2. Formally verify the correctness of every algorithm submitted.
This is as crucial as large-scale deployment for many scientific computing users, and it is one of the banes of (and reasons why) implementing the algorithm yourself. Here your product then could really offer a compelling proposition to these users.
This would also be beneficial for ensuring the reliability of your API, even if you formally waive liability to the algorithm developers (as you surely do). Else you might find yourself on the other side of securities regulators for a multi-million dollar trading glitch caused by one of your algorithms [4], or something crazy like that.
[1] In fact, in the wikipedia article for Numerical Recipes, it is claimed that one of the motivations for the development of the GNU C library was precisely to come up with a free alternative to them! See: http://en.wikipedia.org/wiki/Numerical_Recipes
[2] http://aufbix.org/~bolek/download/nr.pdf
[3] http://www.astro.umd.edu/~bjw/software/boycottnr.html
[4] http://www.bloomberg.com/news/2013-10-16/knight-capital-agre...
cheers