HNHacker News
TopNewBestAskShowJobs

tijsvd

272 karma · joined September 23, 2019

submissionscomments
tijsvd··on AWS engineer reports PostgreSQL perf halved by Linux 7.0, fix may not be easy
From what I understand in the follow up: postgres uses shared memory for buffers. This shared memory is read by a new connection while locked.

In postgres, connections are handled with a process fork, not a new thread. If such a fork first reads memory, even if it already exists, that causes a minor page fault, which goes back to the kernel so it can update memory mapping tables.

The operation under lock is only a few instructions, but if it takes longer than expected, then that causes lock contention. Regression in the kernel handling minor faults?

The whole thing is then made worse because it's a spinlock, causing all waiting processes to contend over the cpus which adds to kernel processing.

Mitigated by using huge pages, which dramatically reduces the number of mapping entries and faults. I reckon that it could also be mitigated in postgres by pre-faulting all shared memory early?

tijsvd··on Where did the false "equal transit-time" explanation of lift originate from?
TLDR is that the full correct explanation is not simple. The Wikipedia article on lift makes a good effort.

https://en.m.wikipedia.org/wiki/Lift_(force)

tijsvd··on Jiff: Datetime library for Rust
Of course you don't need a calendar library to measure 30 seconds. That's not the use case.

Try adding one year to a timestamp because you're tracking someone's birthday. Or add one week because of running a backup schedule.

tijsvd··on The Rust I wanted had no future
Everything that's visible to the compiler is subject to automatic inlining. That is all code in the current crate (compilation unit), all concrete instantiations of generics (regardless where defined), and all functions marked inline (regardless which crate).

Stdlib containers are all in the generics category.

tijsvd··on How to be a -10x Engineer
> Get this, say a project takes a year to complete. The concept is saying a 10x engineer can do this in about month.

The true 10x engineer looks at the project, sees the inherent needless complexity, goes back to the sponsor and uses his business knowledge to renegotiate the specs. Leading to a reduced scope with 98% of the business value and 10% of the work.

tijsvd··on A world to win: WebAssembly for the rest of us
Would it not be feasible to turn electron inside out, and have chromium as a library, with bindings for various languages?
tijsvd··on I love building a startup in Rust but wouldn't pick it again
Never combine new tech with new functionality. If you want to learn new tech, use it to rewrite an old project that was due anyway. If you want to build new functionality, use tech that you know.

This has nothing to do with Rust. I've seen the exact same thing happening with golang in a C++ only environment. Long project, took forever, failed slowly, took a week to rewrite in C++.

tijsvd··on When Rust hurts
You get the same with e.g. tokio::spawn, which runs the future concurrently and returns something that you can await and get the future's output. Or you can forget that something and the future will still run to completion.

Directly awaiting a future gives you more control, in a sense, as you can defer things until they're actually needed.

tijsvd··on Source code for Dutch DigiD app released under Dutch Open Government Act
Then the app relies purely on the ssl cert of the server, for mitm mitigation. This way, the qr can contain a signed reply to the code, which adds a layer.
tijsvd··on We're wasting money by only supporting gzip for raw DNA files
Apart from CGAT being a 2-bit alphabet, changing the alphabet does not change the information density. Expect that this kind of transformation has no impact on the compressed result, with most general purpose algorithms.
tijsvd··on Make your database tables smaller
First, this is about data access and not programming.

Second, that quote is from a time when compiler optimizations did not exist, and the programmer was supposed to use all kinds of clever tricks to speed up code (what today you get for free with -O3). That kind of optimization is the context of the quote, and it's hardly ever appropriate to just throw into some discussion about optimization.

tijsvd··on AI unmasks anonymous chess players, posing privacy risks
It's the field of information theory.

https://en.m.wikipedia.org/wiki/Information_theory

tijsvd··on User IDs probably shouldn't be passed around as ints (2018)
Take incrementing int32. Extend to 64 bits. Multiply by large prime number (e.g. fnv32 prime). Mod 1 million. Add 1 million. End up with random looking, 7 digits, nicely sequenced, 32 bit integers. Write an exhaustive test to verify.

When near 1 million users (yagni), reset sequence and do the same with 10 million (or one billion).

Doesn't solve the upside of 128-bit random numbers (ala uuid): the ability to generate remotely and expect no collision.

tijsvd··on Paginating Requests in APIs (2020)
What do you use then that has the same order as rows becoming visible?

We use an auto-increment id, and lock inserts on the related account (which always limits the scope of the query).

The only other (stateless) way I can think of is to somehow fiddle with transaction numbers linked to commit order.

tijsvd··on Tokio Console
The callback is really hard to implement without allocating memory for each wakeup. The poll mechanism can simply leave the task in place. I suspect the poll thing is also easier to generate.
tijsvd··on Show HN: fcode is a binary rust-serde format that supports schema evolution
I was generally OK with the speed of Prost, and making it faster was not the goal really. The goal was to make the generated code more Rust-friendly.

It is probably faster due to straightforward decoding, no tag switch. At the cost of not being able to reorder fields or remove in the middle.

With protobuf implementations, one always goes from schema to generated code, which means certain choices can not be handled without further markup:

- what to do with unknown enum values

- which fields become Option

- which string or bytes fields should be by-ref rather than by-value

- derive of other traits

I guess the alternative could have been to build all that into Prost, but that seemed unreasonable. I see this more as a replacement of bincode than as a replacement of Prost.

tijsvd··on The Lava Lamps That Help Keep the Internet Secure (2017) [video]
Thanks for that link. The article links to another article with more background:

https://blog.cloudflare.com/lavarand-in-production-the-nitty...

Which answers all your questions. Tl;dr: the lava feed is only an additional entropy source, it's not like they generate keys directly from the video feed.

tijsvd··on A Bit Overcomplicated
Does Rust's `:b` formatter always print leading zeroes? Otherwise the code will not only be inefficient, but wrong for half the input space.

Or perhaps that was the original goal, to find the most significant one and the 41 bits that trail it...

tijsvd··on Why does the New menu even exist for creating new empty files?
IIRC (many years since I used windows), the New context menu on windows is fed from a directory with templates, and you can add to it.
tijsvd··on A Concrete Introduction to Probability (2018)
With 2 children, there are 4 configurations of equal probability. The one with 1 boy 1 girl occurs twice. Take away the 2 girl case, then 2 boys is 1 in 3.
tijsvd··on My tutorial and take on C++20 coroutines
Does that mean that every nested coroutine (async call) needs another heap allocation, or just the top level one?
tijsvd··on My tutorial and take on C++20 coroutines
> And from what I've read they are better than Rusts coroutines for this use case

Reference please? In what sense are they better, and what makes them better?

tijsvd··on Again on 0-based vs. 1-based indexing
Nah, it doesn't matter if you do (prev + 1) % len, or (prev % len) + 1.

Hash tables matter more though.

tijsvd··on Don't Use Protobuf for Telemetry
No I don't work for Blizzard. When I did this, we were using a very fast custom wire format, but entirely hand-coded. Say flatbuffers without the code generation part. And having massive trouble with schema evolution.

Another group had already evaluated pb and found it way too slow. They had designed something similar but faster. I wrote my implementation of pb to prevent this, and showed pb could be fast enough. It was definitely the right choice at the time, as it gave us C++ speed close to the old format, plus easy interop with other languages.

> If protobuf works for Google then it essentially works for 99.999% of every other company on the globe.

Uh.. no. Google is a massive company, but if you browse the comments in this thread, you'll find multiple remarks like "this was built for Google's servers". Google have specific use cases, and they build software for that. The software may well be lacking for other use cases. I can totally imagine Blizzard wanting to write their own implementation, think of the benefits of reducing parse time in a multiplayer server.

tijsvd··on Don't Use Protobuf for Telemetry
Yes, this comes back to use case.

So it's great that there are different implementations for different use cases. It helps that the wire format is simple and well-documented.

tijsvd··on Don't Use Protobuf for Telemetry
Sure that makes some sense.

But it's somewhat ugly as well, from an architecture point of view. Because that argument translates to anything else you might want to do with these objects.

In past libraries for C++ I've tried to prevent this kind of coupling by adding generated template methods like "walk(f)" where f would be a templated callable, called with a descriptor and data reference for each field. Any kind of pretty printer or SQL statement can be built that way.

tijsvd··on Don't Use Protobuf for Telemetry
> No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications.

That's a bit harsh. Protobufs deliver smaller wire size than any of the newer "zero-copy" formats. And many receivers of zero-copy formats will... copy the data into some internal representation. If your protobuf implementation delivers classes that are good enough to work with internally (store in maps, forward, etc) then you don't really lose something; instead you gain, due to no manual conversion layer.

tijsvd··on Don't Use Protobuf for Telemetry
It didn't exist yet. We were evolving from simple raw messages with all sorts of problems.

The alternative was something custom again, with better support for schema evolution, but pb was convenient due to existing implementations in Python (system tests) and C# (UI).

tijsvd··on Don't Use Protobuf for Telemetry
But then all that must come with bookkeeping, which brings its own cost.

Take a look at an implementation like Prost, for Rust. It's very similar to what I did (10 years ago by now). Everything is just inline, except when messages can be recursive (which should be rare for most protocols).

tijsvd··on Don't Use Protobuf for Telemetry
What is this thing with JSON support? Don't people use pb so they do not have to deal with JSON? I'd expect that for a truly lean pb implementation, adding JSON is a 300% increase in code size?
Page 1 of 3Next →