HNHacker News
TopNewBestAskShowJobs

giovannibonetti

949 karma · joined November 6, 2014

Software engineer with lots os experience building web apps and getting the most out of Postgres. Nowadays I mostly work with Docker containers, data pipelines (batch and stream), GCP and AWS. I'm a fan of functional programming and strongly typed languages.

My HN handle is also my Github handle.

submissionscomments
giovannibonetti··on Homa: The end of TCP for AI clusters [video]
I remember listening to Jane Street’s Ron Minsky on their podcast talking about this a few months ago, how TCP becomes the bottleneck in AI clusters. As an electrical engineer, I remember that circuit switching gave away to packet switching due to very sparse usage of the network when there are many actors going through it. It is not very efficient, but that wide variety of traffic makes it hard to optimize it since the flow patterns are too dynamic. A good analogy with car traffic is that downtown there are so many cars going to a large variety of places, that traffic lights – as inefficient as they are – are a solution that at least works good enough.

On the other hand, if the traffic follows a very predictable pattern, a custom implementation can be much more efficient. Specially nowadays machine learning can find much better solutions through reinforcement learning. And AI cluster data flow is much more predictable than what goes over the internet as a whole.

giovannibonetti··on Coding Is Not Solved
> Most software that requires hiring and paying software engineers has low risk tolerance:

I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.

giovannibonetti··on We're gonna need a lot more mathematicians
There is a huge amount of pending work in health research that is nowadays ignored because smart people have better paying jobs available.

Once many of those jobs get automated, I bet there will be many more people working in health research, which hopefully should lead to better health outcomes for society as a whole.

giovannibonetti··on Deterministic Core, Non-Deterministic Shell
For frontend web development, Elm enforces a pure core (pure > deterministic).
giovannibonetti··on Deterministic Core, Non-Deterministic Shell
I have been trying to tell the folks at work the same thing. Ideally, using a language that goes in that direction. For frontend web development, for example, there is Elm, which is pure and functional.
giovannibonetti··on Learn Programming with OCaml
Maybe learn Rust as a middle ground between ML and C, then? It has sum types like ML and teaches you memory management like C.
giovannibonetti··on There's no reason for software to be slow anymore
Shotout to PowerSync for enabling companies to go in the opposite direction and build offline-first apps. My company is a (production) customer, and we recommend it.
giovannibonetti··on The Amazon tax
> Companies want everything to be push instead of pull.

This reminds me of Instagram's search page, which is full of animated videos jumping at you to dissuade you from completing the search you wanted and just click on what they want you to see.

giovannibonetti··on A Preview of DuckDB v2.0
Disappointed, since I was expecting they would rewrite the implementation from C++ to Zig. I bet that would increase the number of positive pull requests they get, since most developers prefer to stay away from C++ nowadays.
giovannibonetti··on Does anyone run Postgres without PgBouncer?
I see many comments comparing PgBouncer with application connection poolers without addressing the conceptual difference between them. Here it is:

1. Most application connection poolers follow a first-in-first-out (FIFO) algorithm, which is simple enough to implement and is enough to make sure the application always has a connection available to connect to the database. It optimizes low latency, and works great from the point of view from the application. The problem is that it has few mechanisms to remove redundant connections, since the application is constantly keeping them all "warm".

2. PgBouncer and very few external poolers follow the inverse idea – last-in-first-out (LIFO), and they optimize for reducing the number of connections that reach Postgres, thus improving its throughput. The idea might seem crazy at first – the last connection used is the first one to be picked up again – but this algorithm automatically removes excess connections, which will get cold and get closed.

When starting a new application, option (1) is enough, but as it scales up enough, at some time it is recommended to use (2), since having hundreds of open connections to Postgres is bad for performance if you can use PgBouncer or similar to cut it by 90%. Postgres' process-per-connection design works much better when there are fewer connections reaching it.

giovannibonetti··on "Solving a largely imaginary user goal"
Not necessarily. You design your data model to prioritize (1) correctness and (2) performance. It doesn't have to resemble the UI at all, as long as the UI can fit on top of it with some abstractions.

Some examples that come to mind: - video games with their entity-component systems; - high-performance text editors like VS Code. [1]

[1] https://code.visualstudio.com/blogs/2018/03/23/text-buffer-r...

giovannibonetti··on Google is making private AI practical with homomorphic encryption
E2E encryption means that if the user loses the keys, there is no way to recover that even if they contact support and prove the data belongs to them.
giovannibonetti··on Accelerating GPT-5.6 Sol Ultrafast
In some romantic languages less and fewer are the same word, so it is a common mistake for people that have them as their mother tongue.
giovannibonetti··on We replaced Redis with MySQL for inventory reservations and it scaled
I wonder how we could handle that in a simpler way with durable workflows (e.g. Temporal, Restante, DBOS) – which are similar to Erlang processes but with persistent disk storage. This could avoid the need to maintain the 1000 row inventory.

Perhaps each shopping cart would have its own workflow, and the inventory item would have one as well. Then, whenever a customer put an item in their cart, their cart workflow would send a signal to the inventory item workflow and wait for the response. The inventory item workflow would maintain a ledger controlling to which cart each unit goes, and it could batch the writes to this table. This way, even if 100k customers try to purchase the same item in the same second, it should handle the load.

After the batch is written to the ledger, the inventory item workflow would reply signals to each cart workflow confirming that the reservation was completed. The end-to-end latency from the consumer point of view would be a fraction of a second, without needing the 1000-row hot-inventory heuristic.

giovannibonetti··on Indexing the Data Lake for Online Point Queries
Instead of spending all of this effort trying to optimize an OLAP data lake for an OLTP use case, why not ust use a regular OLTP database like Postgres or MySQL?

Even if the data size is "infinite" you can put it on PlanetScale or something like that. If the tables are well optimized with covering indexes, you can pull out thousands of rows in 100ms for point queries, which is plenty for "point" queries.

giovannibonetti··on The Open-Source Release of ML Video Codec (MLVC)
> For example, for 360p video at 30 fps, where H.264 requires 1 Mbps, MLVC requires roughly 122 kbps for equivalent quality — about one-eight the bitrate under real-time conditions. The inference compute was kept approximately equal for the 360p and 540p resolutions. These results are based on a P.910 subjective test and are based on the Video Conferencing Dataset (VCD) dataset that we developed and also recently released as an open-source project.

Very impressive. Congratulations!

giovannibonetti··on Private healthcare makes industries less innovative. It's time for change
> It is just a really awful system overall.

You mean for the user. It is great for the insurance companies, though, since tying health insurance to employment indirectly keeps the average user age lower, thus making the overall scheme more profitable.

As a general rule, people use health benefits more as they age.

giovannibonetti··on The startup's Postgres survival guide
> FOR UPDATE SKIP LOCKED > The best way to think about this Postgres feature is that it reserves the rows that you’re selecting for use in your transaction without interfering with other queries. We use it primarily for implementing our job queue;

SKIP LOCKED is useful for implementing job queues with interactive transactions – you lock the row while working on it in the application and keeping the transaction open. For high-performance applications it is best to avoid interactive transactions at all, and just update the rows to "pending" immediately. There is no need for SKIP LOCKED in this case.

As a rule of thumb, as you scale up the application, you want to have less state in the database memory, and interactive transactions are just that. Idempotence beats atomicity at scale.

giovannibonetti··on The startup's Postgres survival guide
> The reason I think it’s useful to view queries as binary—they either seq scan or they don’t seq scan—is: the more you micro-optimize a query, the more of a risk you take that the query planner goes rogue. If you stick to querying by primary keys and indexes, the query planner will have a much easier time.

It's also important to notice the query planner optimizes for the average case, but often it would be better for the app developer if it was optimized for the worst case. But optimizing for the former is a much more tractable problem, so no wonder that's what is implemented.

I had to fight against the query planner when it would optimize a query for the average user, with few rows in a given table, and it would pick one index that made sense for that situation and return a result in less than 10ms. However, when a heavy user issued the same query, depending on the exact parameters the worst case could take over 1 second. So I had to write a much more complex query to force it to take another path with a different index, which would be slower in the average case, but in the worst case would take still less than 100ms. Avoiding timeouts was much more important for my company than taking 10ms more in the average case.

giovannibonetti··on The startup's Postgres survival guide
> Because of all these connection footguns, external connection poolers like pgbouncer are great! If you can’t add this for whatever reason, in-memory connection poolers are a great second option. For example, because Hatchet is open-source, we don’t assume that all user databases use connection poolers, so we use pgxpool (an in-memory connection pool for Go) for this purpose.

Few people know that there is a major bifurcation when it comes to connection pooling implementation.

1. Most application connection poolers follow a first-in-first-out (FIFO) algorithm, which is simple enough to implement and is enough to make sure the application always has a connection available to connect to the database. It optimizes low latency, and works great from the point of view from the application. The problem is that it has few mechanisms to remove redundant connections, since the application is constantly keeping them all "warm".

2. PgBouncer and very few external poolers follow the inverse idea – last-in-first-out (LIFO), and they optimize for reducing the number of connections that reach Postgres, thus improving its throughput. The idea might seem crazy at first – the last connection used is the first one to be picked up again – but this algorithm automatically removes excess connections, which will get cold and get closed.

When starting a new application, option (1) is enough, but as it scales up enough, at some time it is recommended to use (2), since having hundreds of open connections to Postgres is bad for performance if you can use PgBouncer or similar to cut it by 90%. Postgres' process-per-connection design works much better when there are fewer connections reaching it.

giovannibonetti··on Show HN: Jacquard, a programming language for AI-written, human-reviewed code
This reminds me of the Flix programming language [1], which has annotated effects like so:

def main(): Unit \ { Clock, Http, Logger, IO } = ...

[1] https://flix.dev/

giovannibonetti··on DSLs Enable Reliable Use of LLMs
This reminds me of this Bjarne Stroustrup's Rule (creator of C++): - For new features, people insist on loud, explicit syntax. - For established features, people want terse notation

Hillel Wayne [1] argues that the same applies for the differences between what beginners and experts desire from a language: Beginners need explicit syntax, experts want terse syntax.

In my mind, DSLs are related to that – a short notation to avoid repetition. And LLMs are the experts.

I wonder if Lisp with its powerful DSL-creating macros will enjoy more popularity in the near future.

[1] https://buttondown.com/hillelwayne/archive/stroustrups-rule/

giovannibonetti··on Online vs. Offline AI Evals: When to Use Each
> That's fewer moving parts, one load-bearing system instead of two...

When I saw "load-bearing" I remembered the other discussion in the frontpage explaining how to prevent Claude from saying that so often.

giovannibonetti··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
Perhaps what is missing is a better memory/caching layer to avoid doing the same for explorations over and over again.
giovannibonetti··on Protobuf-py: Protobuf for Python, without compromises
I wish LaunchDarkly and other feature flag providers supported protocol buffers to MN define the feature flag schema. It would be a game changer when you have complex variations and end up reaching for untyped JSON.
giovannibonetti··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Property testing and deterministic simulation seem like good alternatives.
giovannibonetti··on Fast Software, the Best Software (2019)
Shout-out to PowerSync for making it easier to develop fast offline-first mobile apps. It pushes data from Postgres/MySQL/SQL Server subscriptions to a SQLite into the user's mobile device, avoiding the need for many loading animations when the data is there ahead of time. My company is a customer and we recommend it.
giovannibonetti··on Show HN: Y – A malleable coding-agent desktop app built with Electron
It would be even cooler if it was made with a Lisp and took advantage of it being homoiconic.
giovannibonetti··on Show HN: Bun-sqlgen – Type-safe raw SQL for Bun, no ORM
Those looking for a more mature solution in this space will probably enjoy SQLc [1]. It was initially developed for Go applications, but over the years it got pluggins for many other languages, including JavaScript/Typescript.

[1] https://sqlc.dev/

giovannibonetti··on Phoenix LiveView 1.2
> but seeing how Lustre does HTML templating versus how Phoenix does Heex was my deciding factor to try the latter. My understanding is this is because of a current lack of any macro system.

When I worked with Ruby on Rails I was "addicted" to macros for everything, but after working for a while with statically-typed languages like Elm and Gleam, I see that there are many other ways to solve those same problems. The code can be quite repetitive sometimes, but as long as the compiler ensures everything is in place, it works quite well.

Page 1 of 16Next →