HNHacker News
TopNewBestAskShowJobs

ahachete

2,486 karma · joined June 4, 2014

https://aht.es https://twitter.com/ahachete
submissionscomments
ahachete··on The product I worked on for the last 4 years is now open-source
> Elastic license doesn't impose restrictions.

It does. It explicitly forbids usage. FOSS doesn't impose any restrictions on usage.

Fundamentally different things. I encourage you to review the four freedoms of Free Software [1] and see how AGPLv3 provide them all while Elastic License does not.

[1] https://www.gnu.org/philosophy/free-sw.en.html

ahachete··on The product I worked on for the last 4 years is now open-source
"quite permissive" and "open source" are terms quite far apart.

Open Source software provides certain clear guarantees to the users of that software. Even small changes to these guarantees probably render the software not open source.

This is not pedantic rhetoric, there are clear reasons why open source software should be clearly told apart from proprietary (including source available): given the OSS guarantees, potential users of a given software may make usage / no usage decisions without further due diligence. If those guarantees are modified, due diligence and risk studies may be needed, specially for companies (what if we're not a competitor today but tomorrow we want to? what do you call competitor? etc).

Open Source exists for a reason, which is to provide a firm ground on those who are good with the guarantees it provides.

Please don't try to blur the line.

ahachete··on The product I worked on for the last 4 years is now open-source
I agree with your sentiment, but I'd like to clarify the following, IMO:

> It's a different restriction.

AGPL doesn't impose restrictions. It provides guarantees (that modified versions will remain AGPL and therefore open source for everybody).

> AGPL lets you do anything as long as it is just as Free (as in freedom)

By the very definition of open source software, you can do pretty much what you want, for your own usage. If you want to distribute (e.g. provide a service) with a modified version, then you need to guarantee that modified version retains the right that the original version granted.

ahachete··on Alpine Linux in the Browser (2020)
Wait until the emulator can run a modern kernel with eBPF support. Then use eBPF to intercept the SSL calls pre encryption, push via a map the requests to userspace, and then translate them to HTTP as you wish. Simple! ;P
ahachete··on Hydra – the fastest Postgres for analytics [benchmarks]
But something that is variable in time cannot be "consistent" by definition ;)

Great to know this is known. My recommendation still holds: publish results with GP3: whatever others do (potentially, wrong) shouldn't prevent you from doing it right.

I'd be giving a deeper look at the project.

ahachete··on Hydra – the fastest Postgres for analytics [benchmarks]
That's an assumption. My #1 rule for a benchmark is that it should have reproducible results. Using variable-performance storage goes directly against reproducible results.

On a related topic: an OLAP benchmark with a small dataset that fits in memory caters only to what I'd consider a small set of OLAP use cases. I'd love to see one with a large dataset much bigger than memory.

ahachete··on Hydra – the fastest Postgres for analytics [benchmarks]
I'm a deep Postgres person and I'm very interested in the project. Congratulations and welcome as a new open source project to the ecosystem.

One important recommendation: please do not run benchmarks on variable-performance storage (gp2 in this case), as it obviously may yield different performance depending on the credit situation, potentially delivering more or less performance to different benchmark runs / scenarios.

Specifically, for 500GB you get the max throughput (250MB/s, which BTW is pretty low for an OLAP-style bench, in my opinion) but only 1.5K IOPS. See [1] for more information.

An easy alternative would have been gp3 volumes, which do not burst performance, and can be set to provide 16K IOPS and 1Gbps.

[1] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/general-...

ahachete··on Garage: An open-source distributed object storage service
> The legal advice to be skeptical of the AGPL is absolutely right.

The legal risk involved by using AGPL software for a company is exactly zero.

AGPL is an open source license which, by the very definition of open source, means that you can freely use the software. Full stop.

The only arguable risk is when modifying the software and on top of that using it in conjunction with other in-house software. But if you are ready to use a proprietary license, you already refrained from modifying the software.

So just use it and end of story. AGPL is a perfectly fine, open source license.

ahachete··on Why we moved from AWS RDS to Postgres in Kubernetes
In case you are interested, I blogged about it last year: https://thenewstack.io/kubernetes-will-revolutionize-enterpr...

TL;DR performance impact should be negligible, could be even slightly negative compared to a VM (when running K8s on bare metal).

ahachete··on Why we moved from AWS RDS to Postgres in Kubernetes
Light travels much slower (~1.5x slower) on a fiber optic, due to the refractive index (~ 1.5) of the fiber.
ahachete··on Learn Postgres at the Playground – Postgres compiled to WASM running in browser
This is a fantastic idea, truly innovating what you can do with Postgres. Kudos!

Postgres, possibly surprising to many, is very "simple": it has essentially no dependencies other than a few OS system calls (open, read, write files) and some optional dependencies (e.g. libssl). Therefore, it is very portable and "easy" to compile on many environments. This includes new environments or ideas like compiling it to WASM.

But you need to come up with the idea. This is a great one and opens the door to other use cases. I hope this serves to push the mindset that Postgres can also be used in lighter-weight environments where SQLite (another fantastic database, don't get me wrong) is often considered as the only viable choice.

edit: typo

ahachete··on Show HN: Pg_jsonschema – A Postgres extension for JSON validation
Shameless plug (StackGres team member here) but StackGres possibly has the largest selection of ready-to-use Postgres extensions [1].

Give it a quick try on any Kubernetes cluster, like k3s on your laptop (one command install), and install any extension from the Web Console or a 1-line in the SGCluster yaml.

[1]: https://stackgres.io/extensions/

ahachete··on Pgo: The Postgres operator from crunchy data
If building a list of alternatives, let me do a shameless plug for StackGres [1], the Postgres platform for Kubernetes with a fully featured Web Console, AMD64 and ARM64 support and more than a hundred available Postgres extensions [2]. Fully open source, no usage restrictions.

[1] https://stackgres.io/

[2] https://stackgres.io/extensions/

Disclaimer: founder of the project.

ahachete··on Custom SQL functions for data analytics in PostgreSQL
I agree.

If the extension availability is the main concern, I'd recommend the open source StackGres [1] operator, which has, possibly, the largest Postgres extension catalog [2] available.

[1] https://stackgres.io [2] https://stackgres.io/extensions/

Disclaimer: founder of OnGres, the company behind StackGres

ahachete··on Make enterprise features open source
> I didn't say it was an 'announcement of an announcement' just that this

But it appears you didn't consider the following part of my comment, where I explained that a) I didn't have any way to know there would be a later announcement; and b) that there's enough interesting information with the current news that is worthwhile to many (definitely was for me), so there's no reason for holding it.

ahachete··on Make enterprise features open source
I don't see this an "announcement of an announcement", as there's concrete and detailed content about the announcement itself. You can read the commit message and even the source code.

Moreover, I'm not affiliated with Citus and I don't know if they are planning later on an official announcement or not.

I stand by the idea of sharing these good news at the earliest for everyone interested to know, look at the code, and plan ahead of time if they need/want to.

ahachete··on Make enterprise features open source
I submitted the news, sorry if I spoiled the announcement! ;) But as soon as I received the news, and these are definitely good news, I wanted to share it with the HN crowd :)

BTW, is item #8 (or any other, for that matter) an alternative to having to use .pgpass? Because this is a notable itch towards automation of Citus, would be great to see this resolved in the open source version.

Congratulations for the commit!

ahachete··on It’s fine, Rewind: Revert a migration without losing data
Does it support migrations that affect more than one table?
ahachete··on What's New in ClickHouse 21.12
Please also note that there are two concepts here that can easily get mixed. For clarification:

* Zookeeper's wire protocol is emulated in CH-Keeper. Nice! So all clients are compatible, etc.

* Zookeeper uses a distributed consensus algorithm called ZAB. Which is not Paxos --but many believes so. CH-Keeper uses Raft, and it can do so as the consensus algorithm is not exposed directly: it is an internal property hidden behind the API and obviously the wire protocol.

ahachete··on What's New in ClickHouse 21.12
> They rewrite the ZK protocol using multi-paxos within CH itself.

It appears to be implemented with Raft, not Paxos (per https://presentations.clickhouse.com/meetup54/keeper.pdf, slide 21).

ahachete··on FerretDB: A truly open-source MongoDB alternative
Yes, good luck! :)
ahachete··on FerretDB: A truly open-source MongoDB alternative
It was both. There were two separate software based on the same underlying technology:

* ToroDB Stampede[1]: MongoDB replica, converting on-the-fly documents to relational structures. Targeting OLAP, as data normalization made queries from some % faster to 2-3 orders of magnitude faster.

* ToroDB Server[2]: what DocumentDB is or FerretDB is planning to be. It was less developed than Stampede, certainly.

[1]: https://github.com/torodb/stampede/

[2]: https://github.com/torodb/server

(edit: formatting)

ahachete··on FerretDB: A truly open-source MongoDB alternative
ToroDB founder here.

Thank you for mentioning this. Unfortunately, yes, ToroDB is no longer being developed. I still believe it's a fantastic idea, and provides significant value. But when it was being built, 5 years ago, the NoSQL (as in "abandon SQL") state of mind was too strong, and the value proposition was not well understood.

I moved to work on what's always been my passion and preference: Postgres, Postgres, Postgres. For those interested, StackGres[1] is what's now my company's focus.

Things may be different today with ToroDB. The technical foundations and ideas are still there. If there would be significant interest by entities that would like to contribute to its development, it could be considered.

I wish good luck to FerretDB. The task ahead is not easy: MongoDB protocol is very simple, but the API is terribly complex and full of nuances. Getting up and running a simple PoC is very simple. Getting from there to a production quality state with notable compatibility is very hard.

[1]: https://stackgres.io

ahachete··on Greenfield – The In-Browser Wayland Compositor
Hi. I find the project very cool and interesting. Good job!

I wanted to comment on what you state in the FAQ [1] about the license: "_We realize this is quite a restrictive license_".

I think this is not correct, nor a good thing. It's more than fine that you use this license and plan to add commercial licensing. But AGPLv3 is an open source license, which provides all the freedoms of the free software. This, categorizing it as "restrictive" sounds almost the opposite of what the license is.

What the license actually does is to ensure that the software will keep the freedoms guaranteed by its license on many circumstances, whereas those freedoms could be removed on proprietary forks if the license where BSD, Apache2 or similar licenses.

Not advocating for or against any license, but I believe the language used to speak about the license is not correct.

[1]: https://www.greenfield.app/faq

ahachete··on PlanetScale is now generally available
> I am estimating that your database space isn't MySQL, which is just fine of course.

You are absolutely right :) My background is strongly on Postgres, you can see from my profile more information if you want to.

So yes, I apologize if some of my questions are not applying or become to obvious for cases that are MySQL-based. But for the most part, I believe principles of operation are the same.

> [other comments]

As mentioned, thank you very much for the detailed information. This completes the picture that I was looking for. I will definitely go in more detail for some of the links provided.

This principle of operation is not too different from something I proposed to a Postgres project some time ago (https://github.com/cybertec-postgresql/pg_squeeze/issues/18). This tool indeed is conceptually pretty similar. It's a shame that supporting schema changes is not part of their focus at this point. It wouldn't do throttling either, but it shouldn't be a difficult feature to add, I guess.

For other users here that may be interested in the Postgres world, there are two tools that perform similar operation (creating a shadow table and filling it in the background), but are both focused on rewriting the table to avoid bloat, rather than for doing a schema migration:

* pg_repack (https://reorg.github.io/pg_repack/): the most used one, relies on triggers * pg_squeeze: already mentioned, uses logical replication

ahachete··on PlanetScale is now generally available
> Thank you! Please first see my comments to parent

Thank you indeed for the time taken to answer all my comments. Now together with all the information here, I understand how it works, and what the trade-offs are.

If my input serves for anything, I'd strongly recommend to take all the information here and write it in a structured way as part of the documentation. I didn't see there any information as valuable as this one. For me, and possibly many others, knowing this information is required in order to make informed decisions about whether to use this or not; and if so, how and what are the trade-offs (e.g. atomizing the changes such that db changes and code changes are independent, which I agree is in general a good thing, but is something to be clearly aware of).

> Again great point and on our radar. To be honest I previously moved away from caring about the exact cut-over time. We designed gh-ost to do just that: stall cut-over until the engineer/developer is happy to sit at their desk. OVer time, we found it was unnecessary. But absolutely there's use cases for both approaches.

For me it's important as cut-over takes some locks. Sure, for a small amount of time. But these locks may create some problems, so that's why I want to be aware. Most of the time are other DDL changes, which are a non-issue here since you already prevent that. But there could be others related to normal db operation. For example, and this may not apply here but does apply with Postgres, such a lock may queue other locks behind (including read-only queries). And if the cut-over lock is itself blocked by other lock (say an explicit table lock), then everything queues on that table and leads to a lock storm, which in turn may cause effective downtime. That's why when we plan migrations or operations similar as this cutover (for example in Postgres a repack operation, which is essentially rewriting a shadow table, in this case just for the purpose or reducing bloat), we really need to take this into account.

ahachete··on PlanetScale is now generally available
I think this comment answers most of my questions: https://news.ycombinator.com/item?id=29248306 Can you confirm (PS) this is how it works? From what I understand here, there are "shadow servers", replicating from the production traffic.

If so, this is cool. I still see some caveats:

* One already mentioned, the scope of migrations is limited to those where both old and new DDL are compatible with the currently running application. If this is the case, I believe it should be clearly advertised as such.

* Being the migration asynchronous, I lose control of when to deploy changes to the application. Even a hook would go a long way, to trigger this.

* Not knowing exactly then the cut-over process is going to happen is also potentially a problem. I understand the cut-over may involve performance degradation (e.g. higher latency) or even connection loss (may you also confirm PS how it is performed?) during some period of time, possibly small. But still, I may need to plan a small maintenance window. But if this is async, I cannot plan the window appropriately.

Neither of this takes away any merits from the solution.

ahachete··on PlanetScale is now generally available
Thank you for your comments, I appreciate it. I'm still not sold, however. I would like to understand the underlying principles, "how this works". I don't need implementation details (happy if they are shared, though) but more on the main principles of operation. Please see my further comments below:

> Git is very bad at analyzing SQL diffs.

Agreed, nothing against. So PS has built-in a nice SQL diff. Neat! But what this really brings? I mean, it's not that there aren't SQL diff tools, tools to manage DDL migrations. Besides this, why not layer it on top of Git? Many orgs and integration tools already have similar workflows (e.g. approval workflows, issue management tools, CI, etc) and if instead of coming up with a new system it would be a layer on top of the existing ones, it would probably have less friction to use. Just my perspective on this, of course.

> Run concurrently to your production traffic

Can you elaborate? How? Do they run on another servers? Or are they waiting on a queue change waiting to be applied? If they run on different servers, what they run there, since AFAIK the migration is only DDL, there's no data?

> Will automatically throttle when your production traffic gets too high, and in particular taking care not to affect replication lag

Same as above: who will throttle, the migration? But what is the migration? Let's use my example: a column type change requires a table rewrite. So the table rewrite will throttle, i.e. slow down? But where is this table rewrite running, on the main server (apparently not) or on a shadow server (apparently either since migrations have no data)? Actually you mention "when your production traffic gets too high". What is "high", can you quantify? We run customers that do dozens to thousands of transactions per second. Is this high enough? Will their migrations ever run, or will wait for very long periods of time, maybe forever?

> Will run completely lockless throughout the migration

How is this possible? Where the migration is running, then? A shadow table, shadow server... none?

> At cut-over point

What's cut-over? Are groups of servers switched? This is what it sounds to me, and that would explain how it could be lock-less and not affecting production traffic. However, it does not explain how data is synchronized from the production database to the migration branch, nor how it keeps being updated with the real production traffic. This is essentially the crux of me failing to understand how this system works.

In general, I apologize if these are too many questions. But in essence, I feel this all sounds really well, but unless I have a deeper understanding of how the principles work, and they are sound to me, I won't be able to recommend this for production usage, as I know from experience the many caveats migrations have. If they are all solved, hats off, but I would appreciate if from a technical perspective this would be more clearly explained.

Thank you!

ahachete··on PlanetScale is now generally available
I'm a bit confused with the "branching" [0] and "non-blocking schema changes" [1] features. I'm confused as they sound like "the next big thing" and I don't see anything special here. Not saying are bad concepts or ideas, the contrary. But not really useful either. Surely I'm missing something, so I would love to hear from the PlanetScale team here if possible.

I have a strong and long Postgres operational background, so I may be also here with assumptions that might be different in MySQL/Vitess/PlanetScale. My main concerns/questions are:

* I can't imagine testing DDL changes without data. Having data there is so important to understand the change and its impact, that I won't do them without data. And unless I'm mistaken, these branches only contain DDL, no data at all ("data from the main database is not copied to development branches").

* While it sounds neat, has a web UI and a CLI, managing branches of a schema and using CI and approval lifecycle... is something that sounds like I could do, and possibly better (as it is more integrated with tooling and workflows) from Git platforms themselves, isn't it? I could do branches, merges, CI, comments on MRs, approval... I could even easily build a deploy queue ("promote") with a CI. Doesn't sound like too hard.

* I don't understand how the "safeness" and the non-blocking nature of changes are ensured. Many DDL changes will take different amount of locks on rows or tables, which may cause some queuing and even lock storms in the presence of incoming traffic. Without incoming traffic, they may run fine. In other words: the impact of a migration can only be determined in combination with the traffic hitting production. How does PlanetScale do this? How for example is handled the case where a DDL changes the type of a column to another type which causes a table rewrite, which essentially locks the table and prevents concurrent writes?

Again, not saying both concepts are bad. Terminology and methodology may be already an innovation. And surely I'm missing a lot. But other than this, I don't see myself using this (testing migrations without data is a showstopper, and not the only one) and I don't see much of an innovation from a safeness perspective here.

Why this system isn't one where thin clones of the database are created as the branches (e.g. like in Database Lab Engine [2]), where you can play with data too, and then some data synchronization is performed to switch over to the branch once done (is not easy at all, but doable with many precautions)? That would be a significant improvement in the process, IMHO.

[0]: https://docs.planetscale.com/concepts/branching

[1]: https://docs.planetscale.com/concepts/nonblocking-schema-cha...

[2]: https://postgres.ai/products/realistic-test-environments

ahachete··on The LZ4 introduced in PostgreSQL 14 provides faster compression
>Plus maybe some pgtune:

Alternative tuning guide: https://postgresqlco.nf/tuning-guide

(Disclosure: part of the team behind it)

← PreviousPage 4 of 13Next →