It does. It explicitly forbids usage. FOSS doesn't impose any restrictions on usage.
Fundamentally different things. I encourage you to review the four freedoms of Free Software [1] and see how AGPLv3 provide them all while Elastic License does not.
2,486 karma · joined June 4, 2014
It does. It explicitly forbids usage. FOSS doesn't impose any restrictions on usage.
Fundamentally different things. I encourage you to review the four freedoms of Free Software [1] and see how AGPLv3 provide them all while Elastic License does not.
Open Source software provides certain clear guarantees to the users of that software. Even small changes to these guarantees probably render the software not open source.
This is not pedantic rhetoric, there are clear reasons why open source software should be clearly told apart from proprietary (including source available): given the OSS guarantees, potential users of a given software may make usage / no usage decisions without further due diligence. If those guarantees are modified, due diligence and risk studies may be needed, specially for companies (what if we're not a competitor today but tomorrow we want to? what do you call competitor? etc).
Open Source exists for a reason, which is to provide a firm ground on those who are good with the guarantees it provides.
Please don't try to blur the line.
> It's a different restriction.
AGPL doesn't impose restrictions. It provides guarantees (that modified versions will remain AGPL and therefore open source for everybody).
> AGPL lets you do anything as long as it is just as Free (as in freedom)
By the very definition of open source software, you can do pretty much what you want, for your own usage. If you want to distribute (e.g. provide a service) with a modified version, then you need to guarantee that modified version retains the right that the original version granted.
Great to know this is known. My recommendation still holds: publish results with GP3: whatever others do (potentially, wrong) shouldn't prevent you from doing it right.
I'd be giving a deeper look at the project.
On a related topic: an OLAP benchmark with a small dataset that fits in memory caters only to what I'd consider a small set of OLAP use cases. I'd love to see one with a large dataset much bigger than memory.
One important recommendation: please do not run benchmarks on variable-performance storage (gp2 in this case), as it obviously may yield different performance depending on the credit situation, potentially delivering more or less performance to different benchmark runs / scenarios.
Specifically, for 500GB you get the max throughput (250MB/s, which BTW is pretty low for an OLAP-style bench, in my opinion) but only 1.5K IOPS. See [1] for more information.
An easy alternative would have been gp3 volumes, which do not burst performance, and can be set to provide 16K IOPS and 1Gbps.
[1] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/general-...
The legal risk involved by using AGPL software for a company is exactly zero.
AGPL is an open source license which, by the very definition of open source, means that you can freely use the software. Full stop.
The only arguable risk is when modifying the software and on top of that using it in conjunction with other in-house software. But if you are ready to use a proprietary license, you already refrained from modifying the software.
So just use it and end of story. AGPL is a perfectly fine, open source license.
TL;DR performance impact should be negligible, could be even slightly negative compared to a VM (when running K8s on bare metal).
Postgres, possibly surprising to many, is very "simple": it has essentially no dependencies other than a few OS system calls (open, read, write files) and some optional dependencies (e.g. libssl). Therefore, it is very portable and "easy" to compile on many environments. This includes new environments or ideas like compiling it to WASM.
But you need to come up with the idea. This is a great one and opens the door to other use cases. I hope this serves to push the mindset that Postgres can also be used in lighter-weight environments where SQLite (another fantastic database, don't get me wrong) is often considered as the only viable choice.
edit: typo
Give it a quick try on any Kubernetes cluster, like k3s on your laptop (one command install), and install any extension from the Web Console or a 1-line in the SGCluster yaml.
[2] https://stackgres.io/extensions/
Disclaimer: founder of the project.
If the extension availability is the main concern, I'd recommend the open source StackGres [1] operator, which has, possibly, the largest Postgres extension catalog [2] available.
[1] https://stackgres.io [2] https://stackgres.io/extensions/
Disclaimer: founder of OnGres, the company behind StackGres
But it appears you didn't consider the following part of my comment, where I explained that a) I didn't have any way to know there would be a later announcement; and b) that there's enough interesting information with the current news that is worthwhile to many (definitely was for me), so there's no reason for holding it.
Moreover, I'm not affiliated with Citus and I don't know if they are planning later on an official announcement or not.
I stand by the idea of sharing these good news at the earliest for everyone interested to know, look at the code, and plan ahead of time if they need/want to.
BTW, is item #8 (or any other, for that matter) an alternative to having to use .pgpass? Because this is a notable itch towards automation of Citus, would be great to see this resolved in the open source version.
Congratulations for the commit!
* Zookeeper's wire protocol is emulated in CH-Keeper. Nice! So all clients are compatible, etc.
* Zookeeper uses a distributed consensus algorithm called ZAB. Which is not Paxos --but many believes so. CH-Keeper uses Raft, and it can do so as the consensus algorithm is not exposed directly: it is an internal property hidden behind the API and obviously the wire protocol.
It appears to be implemented with Raft, not Paxos (per https://presentations.clickhouse.com/meetup54/keeper.pdf, slide 21).
* ToroDB Stampede[1]: MongoDB replica, converting on-the-fly documents to relational structures. Targeting OLAP, as data normalization made queries from some % faster to 2-3 orders of magnitude faster.
* ToroDB Server[2]: what DocumentDB is or FerretDB is planning to be. It was less developed than Stampede, certainly.
[1]: https://github.com/torodb/stampede/
[2]: https://github.com/torodb/server
(edit: formatting)
Thank you for mentioning this. Unfortunately, yes, ToroDB is no longer being developed. I still believe it's a fantastic idea, and provides significant value. But when it was being built, 5 years ago, the NoSQL (as in "abandon SQL") state of mind was too strong, and the value proposition was not well understood.
I moved to work on what's always been my passion and preference: Postgres, Postgres, Postgres. For those interested, StackGres[1] is what's now my company's focus.
Things may be different today with ToroDB. The technical foundations and ideas are still there. If there would be significant interest by entities that would like to contribute to its development, it could be considered.
I wish good luck to FerretDB. The task ahead is not easy: MongoDB protocol is very simple, but the API is terribly complex and full of nuances. Getting up and running a simple PoC is very simple. Getting from there to a production quality state with notable compatibility is very hard.
[1]: https://stackgres.io
I wanted to comment on what you state in the FAQ [1] about the license: "_We realize this is quite a restrictive license_".
I think this is not correct, nor a good thing. It's more than fine that you use this license and plan to add commercial licensing. But AGPLv3 is an open source license, which provides all the freedoms of the free software. This, categorizing it as "restrictive" sounds almost the opposite of what the license is.
What the license actually does is to ensure that the software will keep the freedoms guaranteed by its license on many circumstances, whereas those freedoms could be removed on proprietary forks if the license where BSD, Apache2 or similar licenses.
Not advocating for or against any license, but I believe the language used to speak about the license is not correct.
You are absolutely right :) My background is strongly on Postgres, you can see from my profile more information if you want to.
So yes, I apologize if some of my questions are not applying or become to obvious for cases that are MySQL-based. But for the most part, I believe principles of operation are the same.
> [other comments]
As mentioned, thank you very much for the detailed information. This completes the picture that I was looking for. I will definitely go in more detail for some of the links provided.
This principle of operation is not too different from something I proposed to a Postgres project some time ago (https://github.com/cybertec-postgresql/pg_squeeze/issues/18). This tool indeed is conceptually pretty similar. It's a shame that supporting schema changes is not part of their focus at this point. It wouldn't do throttling either, but it shouldn't be a difficult feature to add, I guess.
For other users here that may be interested in the Postgres world, there are two tools that perform similar operation (creating a shadow table and filling it in the background), but are both focused on rewriting the table to avoid bloat, rather than for doing a schema migration:
* pg_repack (https://reorg.github.io/pg_repack/): the most used one, relies on triggers * pg_squeeze: already mentioned, uses logical replication
Thank you indeed for the time taken to answer all my comments. Now together with all the information here, I understand how it works, and what the trade-offs are.
If my input serves for anything, I'd strongly recommend to take all the information here and write it in a structured way as part of the documentation. I didn't see there any information as valuable as this one. For me, and possibly many others, knowing this information is required in order to make informed decisions about whether to use this or not; and if so, how and what are the trade-offs (e.g. atomizing the changes such that db changes and code changes are independent, which I agree is in general a good thing, but is something to be clearly aware of).
> Again great point and on our radar. To be honest I previously moved away from caring about the exact cut-over time. We designed gh-ost to do just that: stall cut-over until the engineer/developer is happy to sit at their desk. OVer time, we found it was unnecessary. But absolutely there's use cases for both approaches.
For me it's important as cut-over takes some locks. Sure, for a small amount of time. But these locks may create some problems, so that's why I want to be aware. Most of the time are other DDL changes, which are a non-issue here since you already prevent that. But there could be others related to normal db operation. For example, and this may not apply here but does apply with Postgres, such a lock may queue other locks behind (including read-only queries). And if the cut-over lock is itself blocked by other lock (say an explicit table lock), then everything queues on that table and leads to a lock storm, which in turn may cause effective downtime. That's why when we plan migrations or operations similar as this cutover (for example in Postgres a repack operation, which is essentially rewriting a shadow table, in this case just for the purpose or reducing bloat), we really need to take this into account.
If so, this is cool. I still see some caveats:
* One already mentioned, the scope of migrations is limited to those where both old and new DDL are compatible with the currently running application. If this is the case, I believe it should be clearly advertised as such.
* Being the migration asynchronous, I lose control of when to deploy changes to the application. Even a hook would go a long way, to trigger this.
* Not knowing exactly then the cut-over process is going to happen is also potentially a problem. I understand the cut-over may involve performance degradation (e.g. higher latency) or even connection loss (may you also confirm PS how it is performed?) during some period of time, possibly small. But still, I may need to plan a small maintenance window. But if this is async, I cannot plan the window appropriately.
Neither of this takes away any merits from the solution.
> Git is very bad at analyzing SQL diffs.
Agreed, nothing against. So PS has built-in a nice SQL diff. Neat! But what this really brings? I mean, it's not that there aren't SQL diff tools, tools to manage DDL migrations. Besides this, why not layer it on top of Git? Many orgs and integration tools already have similar workflows (e.g. approval workflows, issue management tools, CI, etc) and if instead of coming up with a new system it would be a layer on top of the existing ones, it would probably have less friction to use. Just my perspective on this, of course.
> Run concurrently to your production traffic
Can you elaborate? How? Do they run on another servers? Or are they waiting on a queue change waiting to be applied? If they run on different servers, what they run there, since AFAIK the migration is only DDL, there's no data?
> Will automatically throttle when your production traffic gets too high, and in particular taking care not to affect replication lag
Same as above: who will throttle, the migration? But what is the migration? Let's use my example: a column type change requires a table rewrite. So the table rewrite will throttle, i.e. slow down? But where is this table rewrite running, on the main server (apparently not) or on a shadow server (apparently either since migrations have no data)? Actually you mention "when your production traffic gets too high". What is "high", can you quantify? We run customers that do dozens to thousands of transactions per second. Is this high enough? Will their migrations ever run, or will wait for very long periods of time, maybe forever?
> Will run completely lockless throughout the migration
How is this possible? Where the migration is running, then? A shadow table, shadow server... none?
> At cut-over point
What's cut-over? Are groups of servers switched? This is what it sounds to me, and that would explain how it could be lock-less and not affecting production traffic. However, it does not explain how data is synchronized from the production database to the migration branch, nor how it keeps being updated with the real production traffic. This is essentially the crux of me failing to understand how this system works.
In general, I apologize if these are too many questions. But in essence, I feel this all sounds really well, but unless I have a deeper understanding of how the principles work, and they are sound to me, I won't be able to recommend this for production usage, as I know from experience the many caveats migrations have. If they are all solved, hats off, but I would appreciate if from a technical perspective this would be more clearly explained.
Thank you!
I have a strong and long Postgres operational background, so I may be also here with assumptions that might be different in MySQL/Vitess/PlanetScale. My main concerns/questions are:
* I can't imagine testing DDL changes without data. Having data there is so important to understand the change and its impact, that I won't do them without data. And unless I'm mistaken, these branches only contain DDL, no data at all ("data from the main database is not copied to development branches").
* While it sounds neat, has a web UI and a CLI, managing branches of a schema and using CI and approval lifecycle... is something that sounds like I could do, and possibly better (as it is more integrated with tooling and workflows) from Git platforms themselves, isn't it? I could do branches, merges, CI, comments on MRs, approval... I could even easily build a deploy queue ("promote") with a CI. Doesn't sound like too hard.
* I don't understand how the "safeness" and the non-blocking nature of changes are ensured. Many DDL changes will take different amount of locks on rows or tables, which may cause some queuing and even lock storms in the presence of incoming traffic. Without incoming traffic, they may run fine. In other words: the impact of a migration can only be determined in combination with the traffic hitting production. How does PlanetScale do this? How for example is handled the case where a DDL changes the type of a column to another type which causes a table rewrite, which essentially locks the table and prevents concurrent writes?
Again, not saying both concepts are bad. Terminology and methodology may be already an innovation. And surely I'm missing a lot. But other than this, I don't see myself using this (testing migrations without data is a showstopper, and not the only one) and I don't see much of an innovation from a safeness perspective here.
Why this system isn't one where thin clones of the database are created as the branches (e.g. like in Database Lab Engine [2]), where you can play with data too, and then some data synchronization is performed to switch over to the branch once done (is not easy at all, but doable with many precautions)? That would be a significant improvement in the process, IMHO.
[0]: https://docs.planetscale.com/concepts/branching
[1]: https://docs.planetscale.com/concepts/nonblocking-schema-cha...
[2]: https://postgres.ai/products/realistic-test-environments
Alternative tuning guide: https://postgresqlco.nf/tuning-guide
(Disclosure: part of the team behind it)