Algorithms as Microservices
blog.algorithmia.com
blog.algorithmia.com
Data is where the value is, because you're the only who has it. It's your competitive advantage. It's your secret stash that no one else can access. And if you're willing to ship it halfway around the world to a microservice provider based on flashy marketing, well, I've got a bridge to sell you before your VC money runs out.
Facebook - good luck having a news feed algorithm with no actual stories and status updates
Google - good luck having a search algorithm with no actual search index
Amazon - good luck recommending products to others if you zero data on what others are recommending.
With all 3, you can at least create a minimum viable algorithm if you have the data.
Hence, companies like Palantir and Pivotal that provide this service. To say that the ability to interpret data is ubiquitous is simply wrong.
Sure, generic sort algorithms come "off-the-shelf", but that is not true for all algorithms. Not all algorithms are public knowledge if they are patented, proprietary, or restricted in use/distribution.
You really need both - algorithms and data, right? And systems of execution for algorithms... and networks to interconnect the systems... and potentially humans to review the algorithm's output... and..... and... and..
I think the key is, algorithms can also represent embedded capabilities (information? knowledge? insight? not sure what to call it) generated from data. I mean, look at a RNN - the values/weights/etc. calculated inside the model are calculated from observations/training on the data. Once you disconnect the algorithm from the data itself, the algorithm alone now has embedded value without the data, potentially enormous.
Or, expert/rule-based systems which contain algorithms explicitly designed to embed human knowledge - those algorithms may not even be trained on actual data, they are simply a simplified expression of codified human knowledge.
If you go one step further and try to make some assumption that the way our brains work is by codifying inputs we receive into algorithmic models... and then using those algorithms to calculate based on new inputs from the environment... we can't keep all the data in our heads, but we can probably keep more algorithms than data.
Maybe I'm being too philosophical, but in my opinion, it's not just about the data.
Philosophically you can reduce them to the same thing, but practically speaking you wouldn't really want to.
E.g. with Google Search it is was perhaps the algo at first but now I'd say it is partly that and partly the high barrier to entry in obtaining all the data (although the data itself is freely available)
For those of us who have no data, algorithms are the most valuable things.
Algorithms are public knowledge, commodities to be written and rewritten at will by common, workaday developers.
I don't think the article is talking about sorting a linked list. Clever and non-trivial algorithms are being researched and discovered everyday both in academia and the industry (take ML as an example). Programs = Algorithms + Data (Structures). You can have a competitive advantage on either or both.That being said I agree that having access to data that no one else has is a massive competitive advantage and with your skepticism about the concept being presented.
Some aren't algorithms (or at least not of any algorithmic note) - for example https://algorithmia.com/algorithms/web/ShareCounts count social shares on facebook, pinterest, twitter, linkedin.
There are a lot of text analysis algorithms, but I find that problematic because in many cases the accuracy of the result on the data you pass it is going to depend on the training corpus. And you don't give it that.
When I look at any reasonably complicated algorithm the price is also quite a lot, if for example you want to use AutoTag https://algorithmia.com/algorithms/nlp/AutoTag to extract tags the price is $116 for 10,000 calls. If I am extracting tags to play around with something personally then I won't hit 10,000 calls quickly, but on the other hand it's my money and I would rather implement it myself so as to not spend the money. If I am extracting tags for a real application running in production 10,000 calls could go quite quickly. I would want to develop the in house capability pretty soon.
Maybe it would be useful to build MVPs that used the functionality and then just slot the algorithm in with a TODO replace this as soon as possible so as not to pay lots message. but my personal predilection would be to just build the functionality I need.
This really also goes for other things, like image analysis etc. etc.
1. common algorithms (sorting, graph algorithms, etc) - are commodity and there are tons of free libraries that have good enough implementations.
2. generic data dependent algorithms with many parameters (classification, NLP, neural networks, etc) - many many parameters to tune and they need to be tuned for each type of dataset. Unless you understand the algorithm, there is no real value in playing around with parameters. But if you understand algorithm then you probably want to implement or modify it to your specific needs (speed, complexity, etc). Only people who just want to play around with data might want to try off the shelve stuff to see how it performs but I don't see much market for serious users. I might be wrong here though.
3. Very specialized algorithms (3D surface reconstruction out of points, antenna analysis, etc) - this includes above algorithm types(common and generic data driven) that were specifically tuned to specific purpose. Some companies sell libraries that do that and those companies would probably benefit from Algorithmia.
There are some commercial desktop applications which do this. Autodesk Fusion is one. Some of the more compute-intensive tasks are sent to servers at Autodesk. If you don't like that, you can buy Autodesk Inventor, which is more expensive and requires a big desktop machine locally. (I've used a 24 CPU machine with Inventor. If you're doing finite element analysis or injection molding simulation or ray tracing rendering, it can be worth it.) Autodesk 123 Catch (3D model from photos) also works that way, which allows running it from a phone app.
Another possibility is that the algorithm is proprietary. This comes up with oil well data analysis. But the data is proprietary too, so everybody has to sign tough agreements.
This article mostly looks like a way to have micropaments (maybe not so micro) for software accessed from servers. If you're running on Linux, it's hard to license software to a server, what with virtual machines, cloud systems, and containers being started and stopped.
Unless you work at Pixar, where render farms have been bandwidth-limited from day one.
For algorithms "microservice" architecture is highly dubious. You may perform point-lookups (given X, categorize Y) but you may also perform a large-scale computation (given millions of X, categorize Y). In that case, moving _code_ around instead of data is much more valuable. Thus you get systems like SQL, Hadoop, Spark where code is shipped around instead of data.
https://en.wikipedia.org/wiki/Dataflow
https://en.wikipedia.org/wiki/Dataflow_programming
https://en.wikipedia.org/wiki/Pipeline_(software)
And similar premises. This isn't a totally new idea. Come up with the minimum functionality for a process. Construct a method of communicating and plumbing these together. Spin them up on demand as needed. Many companies have done research on this concept.
But consider we already outsource algorithms for bitcoin generation and transaction processing. A small amount of data that does require specialized processing. We're seeing it again with voice recognition and image processing APIs. This will continue as we have more data which will encourage more specialized hardware and create situations where algorithms as a service make sense, like Ethereum's smart contracts and 21.co's flat pricing.
In theory it could be very efficient for easy code problems to be solved, leaving man-hours for innovation. Which is what we have all worked towards in open-source. But the flattr/gittip model essentially fails in this regard. Can it be solved by the free market?