Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine
oath.com
oath.com
Really goes to show that engg != business and unless you have a firm business model and growth, an amazing engg team can only get you so far.
I know I'm just stating the obvious, but just putting it out there!
Vespa looks super interesting (more so since I'm in a company that provides ecommerce search APIs as a product) and I'm sure I'll play with it more. Thanks Yahoo! :)
Nadella isn't making a particularly strong case for MS either with his seeming distaste for any product of theirs that isn't either mobile or cloud-based (like the ones where they actually have a monopoly to build off and no competition worth speaking of), which seems to arise from that 'chasing growth numbers' mindset.
(The first sentence was "Windows is a service and updates are a normal part of keeping it running smoothly.")
What are you talking about? She's no longer employed and Yahoo has been sold off to Verizon...
So she did exactly what Google paid her to do. Eliminate a competitor.
How dumb do you have to be to hire a CEO from high in the ranks of your biggest rival?
Loss of direction is what stalls big tech companies 9 times out of 10. If you have 10 decision making actors on the board of directors, you have to effectively do 10 different things that each of that guys want, and all of them with compromises made to accommodate the 9 remaining other things that had to be done simultaneously.
My father frequently says that "a public company is a like a woman with 10 husbands, all trying to make love to her at the same time"
a) the concept of 'tough love' is applicable
b) were engineers allowed and challenged to innovate? or did their work effort primarily consist of navigating red tape and beurocracy to the point that all creativity was crushed?
etc.
see also: dilbert.
If you build crappy tools don't be surprised when people don't find a lot of value in them.
It was also kind of clear that it went too far, that a more experienced hardware design and engineering team would have found (not a cheaper), but a more effective way to handle those same challenges.
What you mean to say is that the Juicero was overbuilt and underengineered. As the saying goes, "anyone can build a bridge, it takes an engineer to barely build it".
Especially so after Yahoo conceded Search.
So one side wants to pour money into getting technical leverage. The other side largely just wanted more effective ways of publishing and monetizing content.
It doesn't really matter which side was right, only that they were at times pulling in wildly different directions in terms of what they believed it was important to invest in.
Somehow, over time, this message was lost, and it was no longer ok to be #2 or #3 in some verticals (like search), and having very healthy profit margins wasn't enough either. The Microsoft search deal was a bad deal, poorly executed, as well: Microsoft couldn't meet the monetization and performance requirements from the start (Microsoft paid out of pocket for a while for this), y! lost control of an important product, and the staffing reductions to search that were supposed to be enabled never came.
Put another way should your purpose be something customer focused?
And/or what do you propose that fits what Yahoo! did (or should have been doing) at any point in its life? Alternatively, provide a clear mission statement than encompasses what any large company with lots of products does -- GE, IBM, Google, etc.
And outside of investment in Alibaba, their involvement in online shopping was marginal. In fact, Yahoo's premium services was always marginal across the board - I spent years trying to push product teams to find services they could potentially charge for in Europe, and from what I can tell my replacements had no luck in that respect either.
Instead Yahoo actually divested a number of services in that respect. Yahoo! Personals was sold to Match.com, and while it may have been co-branded for a while, that's long since lapsed in most markets.
So either Yahoo dramatically failed to follow this up in any reasonable way, or the actual intent was a lot narrower. Or both.
But what does "be a part of" even mean? Have an ad on the page? Be recognised by the user as providing the service? Providing content? Provide the tech? It really says nothing. It doesn't say anything about the purpose either. Is it to sell more ads? To grow the brand? To drive people to premium services?
It's a non-statement.
To me it's a statement that's basically carte blanche for whatever management happens to want right now, but it's certainly not providing focus.
And while I would interpret it as directed towards media, I suspect it did nothing to e.g. clarify to the company whether Yahoo was a media company or a tech company. Maybe intentionally, because there were a lot of people in engineering when I was there who did not want to accept that what mattered to Yahoo was connecting eyeballs to content, and that.
It's a little hyperbolic. A realistic refinement would add family friendly and having a decent amount of marketing spend and/or user time spent.
> And outside of investment in Alibaba, their involvement in online shopping was marginal.
Yahoo! Shopping was a decent price comparison tool (and Kelkoo was an acquisition of the same in Europe), and Yahoo! Stores was enabling a lot of small businesses to sell online. Auctions was closed in 2007 (except in Japan), because it's hard to be the number 2 for auctions. Yahoo! Wallet was supposed to make it easier to buy things all over the web by storing your payment credentials in one (trusted) place. There was a person to person payment thing that I can't recall the name of.
> Yahoo! Personals was sold to Match.com
In 2010 -- way into the period where nobody knew what Yahoo! was trying to do.
> To me it's a statement that's basically carte blanche for whatever management happens to want right now, but it's certainly not providing focus.
Yahoo may have been focused on the directory initially, but at least by 1997 it was not a goal to be focused -- there was never one thing Yahoo did to focus around, other than the brand. Even the first archive.org capture from October 1996 has links to random things.
> And while I would interpret it as directed towards media, I suspect it did nothing to e.g. clarify to the company whether Yahoo was a media company or a tech company.
Does it matter if it is a media company or a tech company? The real answer is Yahoo is an eyeballs company. Premium services I guess were technology revenue, but services revenue was never expected to be a large portion of revenue for the company (although certainly it was for some verticals). Yahoo's tech stack enabled small teams to build vertical sites that competed with much larger teams. Yahoo's business stack enabled those small teams to sign large advertisers and content deals. Yahoo's network of sites enabled the small teams to start with lots of eyeballs, without huge marketing spends. Building compelling verticals gets the eyeballs to come back. Having a great internet search product that the eyeballs saw as a great internet search product would have really helped with keeping eyeballs.
"Over the decades Yahoo! has contributed substantially to the greater good, publishing its own code as open source.
Arguably Yahoo!’s greatest legacy once it is a division of Verizon will be big data, after one of its engineers – Doug Cutting – wrote an open-source implementation of Google’s MapReduce that became Hadoop. What followed was an entire ecosystem of startups and projects crunching data at scale – Cloudera, Hortonworks, MapR to name three in a market some calculate will be worth $50bn by 2020."
Hopefully Verizon's open source legacy will become even bigger and make it easier for these kinds of innovations to happen.
We've had some success a few cities over in Verizon Labs... https://verizon.github.io/
Generally when you disclose something (like the fact that you run the open source process at Yahoo) that's a disclosure, rather a disclaimer.
A disclaimer might be considered the opposite "I don't run the open source process at Yahoo, but...", or, more commonly, "IANAL".
Very optimistic calculation...
Steve Jobs responding to a question about OpenDoc in 1997:
> You’ve got to start with the customer experience and work backwards to the technology.
https://mikecanex.wordpress.com/2011/06/08/wwdc-1997-video-s...
* partial document refeeding (i.e. expedite indexing a new field to 20+ billion documents without refeeding everything and staying online handling 100M+ free text queries a day)
* visual similarity search - check out the tensor ranking features [1] [2]
* online elasticity - add/remove replicas / shards online. A must when it could take weeks+ to re-feed from scratch. This is non-trivial to make work smoothly at scale.
* latency / tail-latency on complex queries. p90 reduction from 3,000 to 30 ms.
This is a major gift to the open-source community of a battle-tested search engine that works reliably without babysitting with very large datasets, and simultaneous high query / high feed volumes. Huge debt of gratitude to the team in Trondheim and Verizon/Oath/Yahoo legal & management teams for making this happen. :+1:[1] http://docs.vespa.ai/documentation/tensor-intro.html [2] http://docs.vespa.ai/documentation/tensor-user-guide.html
$ cloc-git https://github.com/vespa-engine/vespa.git
http://cloc.sourceforge.net v 1.60 T=64.10 s (224.5 files/s, 28276.3 lines/s)
--------------------------------------------------------------------------------
Language files blank comment code
--------------------------------------------------------------------------------
Java 6573 106215 102720 537097
C++ 3209 76542 19178 504855
C/C++ Header 2985 42731 57087 158388
XML 389 705 550 139626
Maven 141 133 244 14096
CMake 450 254 560 8452
Perl 57 1124 762 7649
Bourne Shell 196 1257 734 6918
Scala 95 1685 617 6378
Teamcenter def 234 1474 3490 2468
Lisp 4 231 403 2118
HTML 16 211 29 1950
C 7 288 198 1432
Python 6 132 66 556
Ruby 9 39 9 294
Bourne Again Shell 3 35 12 182
Pig Latin 9 39 52 54
make 2 22 8 39
Ant 1 9 17 36
YAML 1 9 1 22
DTD 2 6 6 10
--------------------------------------------------------------------------------
SUM: 14389 233141 186743 1392620
--------------------------------------------------------------------------------[FLASHBACKS]
To be fair, yinst was the least worst part of the systems I was working on (Yahoo!Europe backend feeds stuff.)
There is some misidentification in your list (our .def files are nothing to do with Teamcenter). We use 2 languages because we Java and C++ have different strengths which make each suitable at different layers of the architecture. The rest is a combination of "for good reasons" and "leftover scraps" :-)
https://github.com/vespa-engine/vespa/tree/master/vespa-hado...
Next up, I would really like to see Sherpa/PNUTS (their NoSQL operational database) and Everest (their petabyte-scale Postgres data warehouse) open sourced :)
I don't quite get the diagram of the Vespa Architecture. Is Vespa a middleware between database engine and query parser? This is what puzzles me.
If so, are there other such middlewares available for ie. PostgresSQL that allow hooking "Query Templating Models" (that is it?) generated via Machine-Learning Models? Is it way more complicated than that, or did they overengineer the problem into a monolith? EDIT: Looking at https://github.com/vespa-engine/vespa it seems that it is overengineered, or maybe it consists of individual micro-components like node.js, hmm more questions :(
Is GraphQL such middleware or lower-level?
Does Vespa replace custom Glue-Code between Backend and Frontend that generates such query-sets for content ranking/positioning?
Or what exactly does Vespa solve? I'm sorry, I've read the article, but can't say, yep that's what it is!
EDIT: How else could you solve what Vespa does using Rust, Go, or C/C++ libraries? A very simple or general direction would be immensely useful to understand Vespa =) The project makes the simultanous impression of an immense engineering feat and at the same time a huge code debt.
It's a datastore in its own right (just like ES), but I imagine that e.g. you wouldn't use it to handle transactions.
This blog post shows how Elasticsearch was used to reindex a 136TB dataset with 36B documents[1], so I'm unsure exactly where except for Google/Yahoo Scale companies Vespa is of use. I would like to understand howto utilize it though without adding an umnanagable complexity.
EDIT: Maybe a Vespa Cloud startup, that abstracts the management and makes "Scalability as a Service" by utilizing other Cloud providers.
--
[1] https://thoughts.t37.net/how-we-reindexed-36-billions-docume...
[2] http://docs.vespa.ai/documentation/vespa-quick-start.html
Anyway, I'm happy that we have more options in this space now.
Let me try myself answering my own question, I hope someone hops in and tells me where I'm wrong or how else to improve :)
1) Get PostgresSQL exntensions via "package manager" pgxnclient
1.1) pg_bouncer - For connetion pooling
1.2) yoke - As a high-availability cluster manager with auto-failover and automated cluster recovery
1.3) prestodb.io - Distributed SQL query engine for pgsql
1.4) pglogical - Logical streaming replication for using a publish/subscribe model
1.5) pg_lambda - To create your own AWS (meta) Lambda
1.6) pg_strom - To offload tasks to the GPU
1.7) zombodb - To utilize full-text searching via indexes backed by Elasticsearch
2) Put all together with pglogical and presto to seperate GPU/CPU intensive tasks.
2.1) "Build Missing Middleware" - To design/fuse a query visually that combines multiple backends
2.1.1) Create a binary data-stream by integrating pg_lambda, pg_strom, presto and zombodb
2.1.2) "Build Missing Middleware" - A tensor processing extension to use ML Model evaluations
2.1.3) "Use Missing Middleware" - For data-processing via Machine-Learning models
2.1.4) "Use Missing Middleware"- To output ML processed results into the database
2.2) Partition these queries using "pg_lambda + middleware" to create accelerated and fused query results
So what's missing to create a Vespa alternative using existing technologies is everything in Point 2) if I'm not mistaken. Torrent based replication isn't exactly neccessary, except at Twitter/Facebook scale, but if you reach that stage you can hire a libtorrent author.It would have following properties: decentralized, distributed, resilient, highly-available, software-defined storage & retrieval system.
According to http://vespa.ai/#featurematrix:
FEATURE VESPA ELASTIC SEARCH RELATIONAL DATABASES
ACID transactions •••
Optimized for analytics ••• ••
Optimized for serving ••• • ••
Scalable ••• •• •
Easy to operate at scale •• •
Text search ••• •• •
Machine learned ranking ••• • 2.1.2) - 2.1.4)
Middleware logic container ••• 1.4)
Live reconfiguration ••• 1.2)
And yet I've to admit that even if the Github repository looks quite chaotic, making an alternative, even using existing technologies would be big feat.Initially I would've chosen PostgresSQL as a base, but the "HA-Layer" is something that shouldn't be decoupled and not a later thought. That's why CAS is a much better form of integration. Also integrating the PostgresSQL Engine into a zfs kernel extension ie. would be a mess. And integrating the database engine into a a distributed p2p algorithm would only add compatability issues an no real advantages.
[1] https://en.wikipedia.org/wiki/Content-addressable_storage#Op...
PS: Clever aquisition by Docker! "Infinit.sh is a content-addressable and decentralized (peer-to-peer) storage platform that was acquired by Docker Inc." And in my eyes one of the best implementations and easiest targets that allow adding a database-layer ontop.
https://github.com/vespa-engine/vespa
I'm really curious how it compares to Lucene/ElasticSearch/ELK, which is currently my tool of choice for (faceted) search and recommendation.
http://vespa.ai/#featurematrix
Next thing I'd really like to know is what existing software it builds on top of.
This article has some more details about the history: https://www.cnbc.com/2017/09/26/yahoo-open-sources-vespa-for...
Disclamer: I work on the Vespa team in Trondheim, Norway.
From the repo, it looks like an absolutely huge, monolithic codebase. (It even bundles its own memory allocator!) Do you know if there are plans to break it up into smaller, more manageable pieces?
While I haven't looked at what's required to deploy this beast, operationally speaking, it sounds it might be daunting to run, and for non-"big data" applications might very well be overkill as an alternative to Elasticsearch.
No plans to break it up into pieces (apart from already consisting of modules). It does one thing, it just happens to be a big thing :-)
If you have a mac of Linux box you can have it up and running in 10 minutes. Multi-node production deployments are no different because Vespa manages the nodes, not you directly.
This may be a difference in mindset, but we generally let apps define the schema, so it's in the developer's realm of responsibility, not an operational concern. We also have apps (currently using Elasticsearch) that manage the schema automatically, derived from a high-level application definition. With ES, the app creates a new index with new mappings, shovels data into the new index, and then activates the new index. But it uses the ES APIs to do this, and restarting anything is not on the table.
I'm guessing (hoping) their lawyers made sure to go over their old agreements with a very fine comb to make sure their license for the software allowed them to open source this.
Don't get me wrong, they've added a lot to it, but there's a lot of code in there that could only have come from their purchase of Overture, who had purchased AllTheWeb from FAST (which was itself purchased by Microsoft).
Similar in purpose to Google's TCMalloc: http://goog-perftools.sourceforge.net/doc/tcmalloc.html
Can this be used as a spark+tensorflow replacement ?
is it used inside Yahoo - because Vespa comes with its own tensor processing engine. I wondered who would use one over the other.
$('pre:contains("yql=")').each((i, el) => { el.innerText = el.innerText.replace(/\+/g, ' ').replace(/yql=(.+%3B)/, (m, p1) => 'yql=' + decodeURIComponent(p1)) })Sustaining throughput over long time is important and often overlooked mentioned in benchmarks.
How is the storage layer designed ? disk format ? Can you extend the layer to support different models such as Property graph ?
1) It attracts developers who will provide fixes, bug reports, documentation, etc. for free (or even draw potential full-time developers for the project).
2) It makes the project look more "trendy" and appeals to developers who will try not to use proprietary software; this is the same move that .NET did and it seems to have worked very well.
Sharing code is fine since we use lots of shared code too. We don't sell code. So if someone wants to use Vespa to make an amazing product and make tons of money, we hope they do. We know that sharing Hadoop helped our competitors, but we also know that the revenue stream comes from ads. So we're glad to share code that makes tech better for everyone. As it turns out, many tech companies feel the same way and openly share code with the industry to help us all get to better tech platforms. In the tech space, it's not about grabbing more of the pie, it's about making a bigger pie. The internet revolution is young and the more we built it in the open, the better it will be for all of us.
This is not clear to me, can you please explain? I'm still stuck at thinking "if you help your competitors then you give up some of your market share".
EDIT: your pie analogy is really nice, I guess it means you grow the whole market by sharing tools like Vespa so you get a smaller slice of the bigger pie. I still don't get how the "revenue stream comes from ads" part relates to everything else.
There's a good give and take in the tech world. Sure, our code has helped Facebook and Google eat our lunch. But we don't blame the tech sharing for that -- since they've contributed quite a bit too. We work rather closely with them on a bunch of projects -- which help us all.
Sure, there are some projects that we'd consider the "secret sauce" that really differentiates us from others. We won't open source those. But a lot of code is there 'cuz we need to move bits around quickly. Sharing that code is not going to make or break a multibillion$ enterprise. It's actually going to help make it better in the long run.
To so make money: our sales people to do that, not the tech people. The sales people are given amazing products, and huge audiences to sell to the advertisers. Whereas a podcast might have a million subscribers, a popular radio program have 10 millions listeners, a TV show getting 50 million viewers, or a wireless company have 150 million subscribers, we have over a billion users -- and the advertisers LOVE that. So we all sell ads, and who ever does that better wins. But as tech folks, we're collaborative. It's not really a new thing, it's very much part of the fabric that has helped the internet evolve.
The point of Joel's article is that "Smart companies try to commoditize their products’ complements."
For what I understand, Yahoo is a media company, and as so it may try to commoditize a natural complement of today's media companies, which is data search.
Smart companies commoditize their products' complements (something that needs to be bought with the product) so that whoever wants buy their product has a large variety of offerings to select from. For instance, MS-DOS's ability to run on any standard PC architecture machine commoditized the PC.
It's not clear how commoditizing data search makes selling media to consumers easier, because consumers of media don't buy data search. In fact they expect it to come for free from the media company.
Might it not have the opposite effect instead? That is, it makes starting up a media company cheaper and allows competitors to spend more money on acquiring media?
It's brilliant, in a way. There are risks, but they are small and mitigated. They may even end up selling support and customizations, or enabling that market. The upside potentials are many and the downsides are few and only risk small impacts.
Hell, you can get RedHat for the low cost of nothing, just by signing up for it. On top of that, they will give you every single last line of code you want. They'll give you all of the code, and do it for free.
Yet, they are a successful for-profit company. They don't even accept financial donations, as far as I know. They aren't the wealthiest company, but they are doing quite well and not suffering financially.
Open source doesn't mean no profit. It just means additional rights for the source code and/or user. (Different licenses prioritize different liberties and have different goals.)
No.
That is, Y! Answers was/is a great idea with terrible results. I'm not sure if it should have been moderated better, or if it should have been marketed better. Hell, maybe it should have had a basic literacy test prior to being allowed to post and answer?
I could actually come up with a few hundred ways to have made it better. We know the question and answer format works. It does on many, many sites. It failed there and, largely, that's because of the users. I'm not sure which was worse, the questions or the answers.
I'd love to be given that project and tasked with improving it. It's a great idea, but horribly implemented. I'd like to fix it because I'm fond of trying the impossible and like fixing broken things.
Improving search for it is absolutely not going to help it. No, that's not going to help in the slightest.
Plus answers was done at a time when we were extremely naive about the wisdom of crowds and open source knowledge. Later efforts learned from those early 2000s failures. I'm pretty sure Joel Spolsky specifically called out Yahoo Answers when he was talking about what Stack Overflow was going to do right in his early announcements.
As for the current Answers site, well, we'll see what happens. I know the PM, a delightful person. I don't know the plan. But if there is a plan to make something useful from it, apply. It apparently makes money (otherwise it would have been killed long ago), and that means there's something to work with. When I started at Yahoo I was hoping to get onto the Groups team for the same reason -- a huge challenge to fix something that could be made cool again.