Jepsen: Dgraph 1.0.2
jepsen.io
jepsen.io
I'd like to thank Kyle for his work on Dgraph. We'll continue to expand Jepsen tests to try and identify more issues and fix the remaining ones.
I'm around if you have questions! Cheers, Manish
Between this and your comment above, it restores a little faith in humanity that there are people out there who honestly want to make their work as good as possible and not paper over cracks.
Yes, but having it published so that everyone sees the actual results without sugar coating or outright ignoring problems is what I meant. I didn't mean that the testing itself was unique, but people funding research and letting the researcher publish it is.
(OK, "unique" is much too strong as that's what much of Jepsen does, but I feel like whenever I'm watching the news and there's a public health study talking about how good or not bad X is, it was often funded by an X industry group and reading more about it find it subject to a lot of issues.)
The Snapshot Isolation bank test bug, in particular, is proving a lot harder to reproduce now -- it takes many hours (even up to 8h) to show a violation; which makes debugging harder obviously. The rest 2-3 are sort of related and easier to fix.
I doubt Jepsen blog post would be updated, but once all the remaining bugs are fixed, we'll write a blog post about it on https://blog.dgraph.io.
We had done a separate testing for Badger before Jepsen, where we had identified and fixed a few issues related to crashes. You can read that blog post here:
https://aphyr.com/posts/317-jepsen-elasticsearch
https://aphyr.com/posts/284-jepsen-mongodb
I understand this analysis was paid for by Dgraph so GIFs and memes aren't appropriate... but I do miss them.
There's another aspect to this: now that I've established a reputation and people take my work seriously, I don't have to be snarky or aggressive to get attention drawn to these issues. And the DB landscape has changed a lot in the last five years! Engineers want to provide formalized safety guarantees. Vendors are taking these failure modes seriously instead of dismissing them as irrelevant! Since I'm getting paid, I can invest more time into making these analyses real collaborations with the vendors. So I see my role now as more about helping people reach their safety goals, and a little less about shaming folks for making systems that didn't live up to their claims.
My other motivations--letting users know how to work with the databases they've chosen, and giving case studies to help other engineers test and improve their own systems, well, those are unchanged. :)
I still reread this and pass it around at least once our twice a year: https://aphyr.com/posts/313-strong-consistency-models
It's dense and takes some time to digest but I consider it required reading for anyone working on or with distributed systems.
like the parent and others here, I've lamented this. I appreciate your considered reply. This callout, in particular, is something I hadn't accounted for, but it makes perfect sense (the general idea that "we're getting serious and our clients are serious now too, so posts have to be serious" is just sort of obvious, and still likely true). Cheers to fighting the good fight!
1. write some code
2. submit to Jepsen
3. if errors, goto 1
The problem with this approach is obvious.
The main issue is that distributed database consistency without trivial, performance-killing locking schemes is too complex to prove when writing or using any trivial, local methods based on e.g. SMT solvers or so.
If something like cockroachdb would be proved-consistent on that level, it would be used for applications currently employing pessimistic locking due to a lack of trust in their database, or scaling vertically without really needing to (there are cases which make horizontal scaling cost-prohibitive due to the dependency chains in the algorithms that can solve them, but they can be replaced most of the time).
There's been a lot of work on both of these problems, but right now, proving concurrent algorithms correct, and proving equivalency of those algorithms to executable code, are very much open research problems.
Glad to see the progress made, and bugs found and fixed, for these Jepsen tests.
As a side note, interesting to see a largely Go-based, real-world project plagued with race conditions. Quite a few bugs read "but it was done in a goroutine" or "but mutable capturing in a goroutine", and even a segfault.
Maybe this will naturally bubble down so that the far more interesting technical discussions can bubble to the top, but I fear this sub-thread will eventually drown out the on-topic discussion.
Reference: https://opensource.org/osd
Your updated statement also says "Dgraph is available under an Apache v2.0 based liberal license", which I think is unfairly using the Apache 2.0 name without clarifying that your license in fact is not open source. The language is still dishonest.
If we had changed the Apache license, we would not have referred to it by that name. But given we have not changed the Apache terms, it seems to us more confusing to arbitrarily refer to those terms by a different name.
My take is that Apache + Commons Clause is more liberal than AGPL, something I can expand on in my blog post -- that's clear from our users, who actually appreciated the switch from AGPL to the current license.
Yes, you have.
> We have added an additional term (Commons Clause) and clearly stated that we are applying the two together.
Adding the Commons Clause changes the terms of the Apache license quite substantially, which is obviously the reason you added it.
https://github.com/dgraph-io/dgraph/issues/2416#issuecomment...
This stuff really sours me on the otherwise fairly reasonable license and its users.
>our responsibility is to choose our license terms and clearly communicate them to those who want to use the software
You are not doing this. Your website reads "Dgraph is available under an Apache v2.0 based liberal license", which is technically true but also very dishonest. Your software does not use the Apache 2.0 license, and users of it should not expect to enjoy the freedoms granted to them by the Apache 2.0 license.
>We have not changed the Apache license terms
This is not true. The Apache license explicitly grants freedoms which the following
>We have added an additional term
changes. Don't wrangle words into passable lies, you're not talking to idiots. Your wording is misleading and dishonest, which are two qualities you should be trying very hard not to associate with your brand. So far you're not doing well. No one is questioning your right to license the software as you wish. But you are not proud of your choice and are using slimy and dishonest language on your marketing material to hide your shame and trick potential users.
A separate issue is the subjective one, which is whether or not the commons clause is liberal. I think you are being dishonest here. There are widely accepted definitions for "liberal" among the open source community. The main differentiating factor of a liberal license is that it's not viral. Most people informed on this subject can then neatly and objectively file licenses into liberal and restrictive categories, with licenses like MIT, BSD, and Apache in the liberal camp, and the GPL family of licenses in the restrictive camp. However, the distinction has always been subjective, and nothing like the OSD exists. Additionally, the distinction is only ever applied in the comparison of open source licenses, of which yours is not a member. The realm of proprietary licenses masquarading as open source licenses is unexplored territory, and I'm drawing a line: no matter how restrictive a license like AGPL is, a license which is not even open source in the first place is far from liberal.
The third issue is whether or not using the Commons Clause is the right thing for Dgraph to do. In my opinion, it's a slap in the face to everyone who ever used or especially to anyone who ever contributed to Dgraph. In fact, as I was researching dgraph a bit more to see how you approached this issue, it also occurs to me that I can't find any discussion where existing Dgraph contributors were consulted on the matter of relicensing their work, which seems to me that your license change was not only distasteful and harmful to your project, but also illegal. Were all of the contributors consulted and their permission obtained to change the license? I'll certainly be consulting them if not.
They degrade on the 'de facto' definition of OS by saying that others may not sell it.
That's it.
Comments like yours are why I don't interact with open source at all.
>The Commons Clause nullifies pretty much any privledge granted by the original license, further rendering it useless. For example, the permission to make and publish changes to the work, or reuse the code elsewhere. Is the Commons Clause viral? If I take a small function from RedisLabs and incorporate it into my project, do my users have to pay RedisLabs to support my project? If RedisLabs decides not to, is my project now illegal to support at all? What if I want to fork the software, does my fork inherit the clause and do the same problems apply? This doesn't sound anything like the "commons" to me. Pretty much all of the rights afforded to users of open source are rendered null and void, and the mention of an open source license in the licensing terms of such software is laughable.
Dgraph was interesting to me until they said they wanted to take a vig for me sharing knowledge and practical advice about the software (because that's what consulting, an activity expressly precluded in the dishonestly-named Commons Clause, is). That? That can go straight to hell.
Being a taker of open source (because you certainly "interact with open source") is your right but has absolutely nothing to do with those of us actually expecting those draping themselves in the banner of open source to practice open source principles and to run an open source project as an open source project. That exists between your ears and your ears alone.
Well, given that there are like 10 different criteria to meet it very obviously is not black and white.
You can view dgraph's source. You can modify and contribute to it. You can fork it, change it, etc. You just can't make it your own product that you sell.
It is open source in many, many ways, and probably the general, colloquial way. Whining endlessly because they don't want other people to sell it, and to maintain some semblance of ownership, is what's ridiculous.
You can not sell the code, or services that are directly based on the code.
Sounds pretty open and reasonable.
But we'll just have to disagree I think.
Wait, did they not have a CLA?
You can use it freely for open source and commercial applications with one, single, easy-to-understand catch:
You can't sell it or provide it as a paid service.
Is it open source? No.
Can I and you and almost everyone except Google Gloud, Amazon AWS and Microsoft Azure use it exactly as if it is open source? Yes.
Is it less hassle than the AGPL? IMO, clearly yes.
No, definitely not. I elaborated on why "liberal" is inappropriate here in a different comment:
>There are widely accepted definitions for "liberal" among the open source community. The main differentiating factor of a liberal license is that it's not viral. Most people informed on this subject can then neatly and objectively file licenses into liberal and restrictive categories, with licenses like MIT, BSD, and Apache in the liberal camp, and the GPL family of licenses in the restrictive camp. However, the distinction has always been subjective, and nothing like the OSD exists. Additionally, the distinction is only ever applied in the comparison of open source licenses, of which yours is not a member. The realm of proprietary licenses masquarading as open source licenses is unexplored territory, and I'm drawing a line: no matter how restrictive a license like AGPL is, a license which is not even open source in the first place is far from liberal.
I also wrote here about how this affects people like you and I:
> You can't sell it or provide it as a paid service.
The limitation in the license text is not that narrow; restricting the right to sell any product or service whose value derives “substantially” from the Commons Clause software. In legal context, “substantially” generally is a rough synonym of “nontrivially”.
You seem to be suggesting that he license intended to only restrict sales of software/services which have no substantial source of value besides the upstream software, but that's not how it's actually written.
As far as I understand the license clearly permits me to:
-Use the licensed software internally.
-Modify the licensed software.
-Share my modifications.
-Write software that connects direcly to the licensed software, with no restrrictions on my software (unlike AGPL)
-From this it follows that I can license my software that uses the licensed software any way I want.
-I'm not sure if I will be allowed to bundle the licensed software with my software,
-But it should be allowed to say that the user as to install the licensed software and run it for my software to work.
-Sharing the licensed software in a bundle (e.g. Linux distro) seems to be OK as long as I don't charge for it. This probably will need to be sorted out for Redhat etc.
Please correct me if any of the above is incorrect. I might have misunderstood it but this was my best understanding.
To me it seems more correct to say:
but you don't get all the benefits you'd have from it being open source.
Again though:
- while I see where this is coming from
- and it won't hurt me directly as far as I understand
- I still have a sketchy feeling about what this will lead to in the future
Yes it is. You can read the source code.
It's not FOSS, or Open Source per the OSI definition.
You can read the source of Oracle Java as well. Doesn't make it open source.
Open source has a defined meaning.
Common Clause even points out in their FAQ that it makes the software not open source.
MIT, BSD, ZLIB, Apache2, etc. are liberal.
That’s the server side version of the anti-tivo clause that made the GPLv3.
I see a v4 coming
I guess the reason I have trouble assuming bad faith with things like this is because it would be much easier for these people to just get high paying jobs at big software companies than to try to eke out a living making a useful source-provided (it would be so much easier to just use "open source" here, but that is verboten) product while arguing with people like you.
Yes, but bad faith exists, and the assumption of good faith is not always appropriate. I don't think that Dgraph is operating in good faith, and the evidence is enough that I am unable to suspend my disbelief. If I were Dgraph and my actions were so severely wrong that people were assuming bad faith, I would want to quickly correct that. But in fact, most of the concerns have gone un-answered.
Well, yes, of you assume the honesty of the person being accused of being deliberately misleading, then that forces you to reject the claim that they are being deliberately misleading, but that is also a prime example of circular reasoning.
I guess I was being too cute in my comment, so I'll switch to frankness. As a total outsider to all this, your and others' assumption of bad faith and willingness to accuse people of outright lies here is very unconvincing and off-putting to me.
I think they're probably just trying to make a living off of software they put a lot of time into creating. It's reasonable to disagree about their approach, but you're uncharitably going further than that.
They claim that Dgraph started with the idea that every startup should be able to have the same level of technology as run by big giants. They started with VC funding in 2016 and their goal was always to make money. That is clearly biggest lie. There is no harm in creating a company for profit but i would appreciate if they were honest.
The only reason Dgraph is open source is to get early adopters and free testers and contribution from community.
It prevents relicensing under the Commons Clause, but I don't see that GPLv3 + Commons Clause would be impossible as an original license, though it has more inherent conflicts to create ambiguity as to actual legal effect than Apache 2.0 + Commons Clause.
>All other non-permissive additional terms are considered “further restrictions” within the meaning of section 10. If the Program as you received it, or any part of it, contains a notice stating that it is governed by this License along with a term that is a further restriction, you may remove that term.
That strictly doesn't prohibit GPLv3 + Commons Clause as an original license, but would seem to make it functionally equivalent to bare GPLv3.
Commons clause means the software works as liberal easy-to-use open source for 99.9% of us AFAIK.
Edit: is->us
Common Clause is — by its own admission — not open source.
I think it is called a strawman.
Edit: as can be seen in my other posts I know it is not open source. As can be seen in their FAQ they know it as well.
Edit 2: attack me -> refute my argument.
> It [sic] this “Open Source”?
> No. [also sic]
> “Open source”, has a specific definition that was written years ago and is by the Open Source Initiative, which approves Open Source licenses. Applying the Commons Clause to an open source project will mean the source code is available, and meets many of the elements of the Open Source Definition, such as free access to source code, freedom to modify, and freedom to re-distribute, but not all of them. So to avoid confusion, it is best not to call Commons Clause software “open source.”
Additionally, "Apache 2.0 + Commons Clause" is not an Apache license. I have seen a number of people be confused by this. I hope Apache implements some guidance with regards to trademarks on this matter, mostly so I don't have to keep saying it.
Anyway we should avoid creating this confusion with this new crazy license, commons clause is malicioua in ita attempt to use a well understood term in a way that twista its meaning.
In software licensing a “permissive license” is a well-established term for a free or open source license that does not restrict downstream derivatives to use the same license (or the license family) for the downstream contributions, as opposed to a copyleft license which does impose such a requirement.
Apache 2.0 is, indeed, a permissive license, and that is a central purpose of its intent.
(Anything) + Commons Clause is not a permissive license, and that too is a central purpose of its intent.
> with one, intentionally very narrow exception.
The restriction in the Common Clause is not “very narrow” (or with very clear boundaries, which makes the area of legal risk larger.)
They clearly intend it to be very narrow it seems.
Can you point out any way it will restrict me or anyone else from using it internally?
BTW: It may seem I think this is all good.
I do not think that.
I do think the cases I've seen so far can be reasonable.
The bigger and more problematic picture I think I see is:
- if every project and its dog applies this that will hurt me if I rely on cloud hosted services
- if the big actors think this is a good idea and start applying thos or something worse to their currently open source software, restricing competition on hosting.
- hollowing out of the open source movement.
Because that's never been how I understood it. I'm pretty sure I understood it to be, at least in it's most encompassing sense, that the source was visible long before the time OSI says it created the label.
Edit: s/weakest/encompassing/
It's deliberately misleading people to ape this terminology to promote software which does not guarantee the fundamental freedoms of open source. If you want to use a proprietary license, own up to it, don't try to get the mindshare of open source when you aren't.
The ancient Greek empire might have included Italy in the past, but an Italian who says they're from Greece today is lying.
That may be, but the terminology chosen by OSI is easily misunderstood, exacerbating this problem. When I see "open source", I think mostly of the source and can I see it, possibly whether I am allowed to alter it for myself. Like an "open book" (hey, I can freely annotate my own copy).
Whether something is free to sell, or redistribute, or alter and redistribute always seemed like variations on what the easily inferred (even if incorrectly inferred) core meaning of open source.
Interestingly, I think the vast majority of people probably think open source means the source is visible, regardless of whether it's an OSI approved license, and regardless of whether they are happy with the license. If that's true, and common usage is at odds with OSI intention, where does that leave us? I think I could easily argue either side in that case.
It leaves us with an education problem.
Ten years ago I don't recall ever seeing people confuse the term Open Source with source-code-available - but the Open Source community was much, much smaller then so I guess the people involved were all on the same page.
Now that Open Source has genuinely won (when's the last time you hard a company say they have a policy of NOT using open source?) it seems we have a new terminology problem that I was previously unaware of.
So what were we referring to Linux and the different software running on it back then as? I was running Linux back in 1996, and I can't seem to remember any other way it was referred to.
Free Software. Sometimes with “free as in speech” to distinguish from the also common “free as in beer” no-cost but restrictively licensed (or, often at the time, with no express license) software.