HNHacker News
TopNewBestAskShowJobs

latch

13,744 karma · joined March 16, 2010

https://www.openmymind.net
submissionscomments
latch··on The Evolution of Caching Libraries in Go
Author of ccache here.

I've barely touched Go in over a decade, but if I did, I'd probably still use ccache if I didn't need cutting edge (because I think the API is simple), but not if I needed something at huge scale.

When I wrote ccache, there were two specific features that we wanted that weren't readily available:

- Javing both a key and a subkey, so that you can delete either by key or key+subkey (what ccache calls LayeredCache).

- Having items cached that other parts of the system also have a long-living reference to, so there's not much point in evicting them (what ccache calls Tracking and is just a separate ARC mechanism that overrides the eviction logic).

It also supports caching based on arbitrary item size (rather than just a count of items), but I don't remember if that was common back then.

I've always thought that this, and a few other smaller features, make it a little bloated. Each cached item carries a lot of information (1). I'm surprised that, in the linked benchmark, the memory usage isn't embarrassing.

I'm not sure that having a singl goroutine do a lot of the heavy-lifting, to minimize locks, is a great idea. It has a lot of drawbacks, and if I was to start over again, I'd really want to benchmark it to see if it's worth it (I suspect that, under heavy write loads, it might perform worse).

The one feature that I do like, that I think most LRU's should implement, is to have a [configurable] # of gets before an item is promoted. This not only reduces the need for locking, it also adds some frequency bias to evictions.

Fun Fact: My goto interview question was to implement a cache. It was always rewarding to see people make the leap from using a single data structure (a dictionary) to using two (dictionary + linked list) to achieve a goal. It's not a way most of us are trained to think of data structures, which I think is a shame.

(1) https://github.com/karlseguin/ccache/blob/master/item.go#L22

latch··on Zig's dot star syntax (value.*)
Author here. I agree this is a less-than captivating piece. I write a lot about Zig and wanted something I could reference from other pieces.

But, to answer your question directly: absolutely. In addition to writing a lot about it, I maintain some popular libraries and lurk in various communities. Let me assure you, beginner memory-related questions come up _all the time_. I'd break them down into three groups:

1 - Young developers who might have a bit of experience in JavaScript or python. Not sure how they're finding their way to Zig. Maybe from HN, maybe to do game development. I think some come to Zig specifically to learn this kind of stuff (I've always believed most programmers should know C. Learning Zig gets you the same fundamentals, and has a lot of QoL stuff).

2 - Hobbyist. Often python developers, often doing embedded stuff. Might be looking to write extensions in Zig (versus having to do it in C).

3 - Old programmers who have been using higher level languages for _decades_ and need a refresher. Hey, that's me!

latch··on In Zig, what's a writer?
There's definitely overhead with the GenericWriter, seeing as it uses the AnyWriter for every call except `write` (1)

    genericWriter        - 31987.66ns per iterations
    appendSlice          - 20112.35ns per iterations
    appendSliceOptimized - 12996.49ns per iterations

`appendSliceOptimized` is implemented using knowledge of the underlying writer, the way that say an interface implementation in Go would be able to. It's a big part of the reason that reading a file in Zig line-by-line can be so much slower than in other languages (2)

(1) https://gist.github.com/karlseguin/1d189f683797b0ee00cdb8186...

(2) https://github.com/ziglang/zig/issues/17985

latch··on In Zig, what's a writer?
A couple comments have mentioned that whatever issue my post raises, are minor compared to the simplicity of the language. But the current implementation of GenericWriter + AnyWriter (with some performance pitfalls) seems more, not less, complicated. Also, neither of these, nor anytype, lend themselves to _simple_ documentation, so that seems like another strike against the simplicity-argument.

As for anytype specifically, in simple cases where it's being used in a single function, you can quickly figure out what it needs.

But in non trivial cases, the parameter can be passed all over the place, including into different packages. For example `std.json.stringify`. Not only does its own usage of `out_stream: anytype` take more than a glance, it passes it into any custom `jsonStringify` function. So you don't just need to know what `std.json.stringify` needs, but also any what any custom serialization needs.

latch··on Fly.io outage – resolved
In most BIG banks, "Vice President" is almost an entry-level title. Easily have 1000s of them. For example, this article points out that Goldman Sachs had ~12K VPs out of more than 30K employees: https://web.archive.org/web/20150311012855/https://www.wsj.c...
latch··on An Analysis of the Performance of WebSockets in Various Programming Languages (2021)
Their explanation for why Go performs badly didn't make any sense to me. I'm not sure if they don't understand how goroutines work, if I don't understand how goroutines work or if I just don't understand their explanation.

Also, in the end, they didn't use the JSON payload. It would have been interesting if they had just written a static string. I'm curious how much of this is really measuring JSON [de]serialization performance.

Finally, it's worth pointing out that WebSocket is a standard. It's possible that some of these implementations follow the standard better than others. For example, WebSocket requires that a text message be valid UTF8. Personally, I think that's a dumb requirement (and in my own websocket server implementation for Zig, I don't enforce this - if the application wants to, it can). But it's completely possible that some implementations enforce this and others don't, and that (along with every other check) could make a difference.

latch··on Putting a full power search engine in Ecto
I don't understand why people use Ecto (or ActiveRecord, or...)

Back in the day, I'm pretty sure we were using Hibernate and friend because our software was shipped and we wanted it to work with whatever database the client was using.

But for a hosted software, what's the point? Not having to know SQL or details about PostgreSQL / the underlying DB ? Apps should be using SQL directly, and for cases where you need dynamic SQL (like, you're where clause is different based on some query string parameters), you can have a low-level query builder (1)

(1) I'm not affiliated with it and have never used it, but a good search came up with https://github.com/robconery/moebius which, at least from the readme, is roughly what I'm talking about.

latch··on Tiered support is an anti-pattern (2013)
In my experience, the real problem with any support scheme is that the decision makers don't have enough skin in the game. It's easy to deprioritize tech debt and bug fixing when you aren't facing 3am pages.

The obvious solution to this is a distributed team. There can still be some holes (e.g. Jan 1 is, afaik, pretty universally celebrated), but having worked at companies that had a team in the US, Europe and Asia was...really nice with respect to this.

latch··on Valkey achieved one million RPS 6 months after forking from Redis
It's a bit like row vs column store.

If there are two lists, in the first example, you're doing:

    a1 -> b1 -> c1 -> d1
    a2 -> b2 -> c2
In the 2nd example, you're doing

    [a1, a2]
    [b1, b2]
    [c1, c2]
    [d1]
Visiting the same number of nodes, but because the nodes are referenced from an array, when you load `a1` you're [probably] also going to load `a2` into the cache.
latch··on Has anyone built their own authentication system?
Why never for a live system?

Store users with an username/email and scrypt-encrypted password.

On login, pull the encrypted password where username = $1. Compare. If valid, create a session id (fill 16 bytes with a cryptographically secure random number generator and encode it), store it that in the db along the user_id and some expiration time.

You now have a session_id -> user_id mapping which can.

latch··on Saving months of compute time with a single Grafana query
Serious question, are you familiar with dedicated hosting?
latch··on Saving months of compute time with a single Grafana query
If your scale is crazy, or your product doesn't allow you to use battle-tested pieces, then orchestration is complex in both cases.

In most cases, managing software on bare metal is more complex in exactly one case: when engineers only know cloud abstractions.

latch··on Leveraging Zig's Allocators
This example came from a real world http server. Admittedly, Zig's "web dev" community is small, but we're trying :) I'm sure a lot could be improved in httpz, but it's filling a gap.
latch··on Leveraging Zig's Allocators
The inspiration for the post came from my httpz library. The fallback using a FixedBufferAllocator + ArenaAllocator is used. The fixed buffer is a thread local. But the arena allocators belong to connections, of which there could be thousands.

You might have 1 fixed buffer, for N (500+) ArenaAllocators (but only being used per one at a time). This allows you to allocate a relatively large fixed buffer since you have relatively few threads.

If you just used retain_with_limit, then you'd either have to have a much smaller retained size, or you'd need a lot more memory.

https://github.com/karlseguin/http.zig/blob/c8b04e3fef5abf32...

latch··on Leveraging Zig's Allocators
fixed, thanks.
latch··on Protecting sensitive data in Elixir GenServers (2023)
No mention of setting the :sensitive process flag to true. Been a while since I've done Elixir, but I think sensitive: true disables additional things, possibly to the point of being a little too much in some cases.
latch··on Zig 0.12.0 Release Notes
I think there are at least 2 serious issues with TLS, both documented in: https://github.com/ziglang/zig/issues/15226
latch··on Zig 0.12.0 Release Notes
Yes. I think I don't understand something, because without a cohesive concurrency story, it isn't clear how you're supposed to glue different components together effectively.

For example, an http server is using nonblocking calls and dispatches to its own threadpool. The lack of cohesiveness means that the blocking call made by the PostgreSQL library can't feedback into the http server's concurrency model. The http server could expose or pass some control mechanism into the application, but for that to work, all the libraries would need to support it.

A lot of people don't seem to think this is necessary, and I've generally assumed that they're right and there's stuff I just don't understand.

latch··on Zig 0.12.0 Release Notes
To add more concrete examples, you could look at tagged unions, errorsets and optionals.

An `if` statement in Zig has a few different faces, traditional condition, unwrapping an optional and unwrapping an errorset.

That makes the language "bigger", but these are things you'll regularly face in C anyways, even with trivial examples. I think most of the things Zig (purely as a language) has "added" to be labeled a "Better C" are all pragmatic and don't increase any cognitive load.

The other thing worth mentioning is that Zig has a number of safety advantages. I just touched up my DuckDB driver, and their C API exposes the underlying storage for a column as a (void ), the format of which depends on the column type. In Zig, I can map this to a typed tagged union. For an bigint column, instead of a (void ), I end up with an []i64.

latch··on Zig 0.12.0 Release Notes
I've written a lot of Zig, including a somewhat popular http framework, a driver for PostgreSQL and one for DuckDB.

Personally, I feel that if you want to grow as a programmer, no matter what you're doing, there's little that offers more bang for your buck than learning C. You can pick up C relatively quickly, and that knowledge will serve you (maybe mostly indirectly) throughout your entire career.

Learning Zig is a much better option. Too many quality of life improvements over C, and you still learn the same fundamentals. If you're interested, I wrote something that might help (1).

As for _using_ Zig. I think it's a good idea to take a conservative view and assume the ecosystem is hostile. The quality of test coverage of the standard library varies. Breaking changes are frequent and aren't always trivial. 3rd party libraries are often abandoned and many are written by people learning Zig (including my own!). It's really great for learning/side/fun project because you can get distracted for weeks writing something you never intended to.

It's not actively hostile, but I think approaching it that way will help manage expectations.

I would strongly consider tracking the master branch. If you write anything substantial in Zig, you will 100% run into a must-have bug-fix or feature in the stdlib or 3rd party library that depends on on 0.13.x development commit.

(1) - https://www.openmymind.net/learning_zig/

latch··on Almost every infrastructure decision I endorse or regret
I don't mean to pick on your specific comments, but I find these analysis almost always lack a crucial perspective: level of knowledge. This is the single biggest factor, and it's the hardest one to be honest about. No one wants to say "RDS is a good choice . . . because I don't know how nor have I ever self managed a database."

If you want a different opportunity cost, get people with different experience. If RDS is objectively expensive, objectively slow, but subjectively easy, change the subject.

latch··on Show HN: Tokamak – Server-side framework for Zig
The biggest issue with std.http.Server is [1]:

   This server assumes *all* clients are well behaved and standard compliant; it can and will deadlock if a client holds a connection open without sending a request.
Atop your readme, you point out that nginx or another reverse proxy should be used. Kudos for that.

As for performance, I'd be curious what gains you get using `std.http.Server` with keepalive and a threadpool. Possibly you can re-use your ThreadContext - having 1 per thread in the threadpool that you can re-using. `std.Thread.Pool` is also very poorly tuned for a large number of small batch jobs, but that's a place to start.

[1] https://github.com/ziglang/zig/blob/b3aed4e2c8b4d48b8b12f606...

latch··on Show HN: Tokamak – Server-side framework for Zig
Like many http server implementations in Zig, this is a wrapper around `std.http.Server`. `std.http.Server` only exists as a mechanism to test the `std.http.Client` (which is needed for things like, downloading packages). So `std.http.Server` isn't built to be fast, secure or robust.

I believe you _can_ get keepalive working with `std.http.Server`, but it seems like Tokamak isn't using it that way. The implementation is a thread-per-connection. Those are the two most obvious issues specific to this implementation.

But I believe the bulk of the issues relate to `std.http.Server`. It shouldn't be public/used.

latch··on Show HN: Tokamak – Server-side framework for Zig
As a Zig developer with one of the more popular HTTP server libraries around (http.zig), this is impressive and good use of comptime and (good) abuse of Zig's anytype. I'm looking forward to learning from the codebase.

But, it's very slow. A quick test shows Sinatra is about 2x faster. Maybe I'm in the minority, but I feel that a primary reason to give up a GC is for performance.

latch··on epoll: The API that powers the modern internet (2022)
In my project, you can start N workers, each accept(2) thanks to REUSEPORT[_LB] and manages its own epoll/kqueue. When a request is complete, the application's handler is called directly by that worker.

I considered what you're suggesting: having a threadpool to dispatch application handlers on. It's obviously better. But you do have to synchronize a little more, especially if you don't trust the client. While dispatched, the client shouldn't be able to send another request, so you need to remove the READ notification for the socket and then once the response is written, re-add it. Seemed a bit tedious considering I'm hoping to throw it all out when async is re-added as a first class citizen to the language.

The main benefit of my half-baked solution is that a slow or misbehaving connection won't slow (or block!) other connections. Application latency is an issue (since the worker can't processed more requests while the application handler is executing), but at least that's not open to an attack.

latch··on epoll: The API that powers the modern internet (2022)
I have a somewhat popular HTTP server library for Zig [1]. It started off as a thread-per-connection (with an optional thread pool), but when it became apparent that async wasn't going to be added back into the language any time soon, I switched to using epoll/kqueue.

Both APIs allow you to associate arbitrary data (a `void *` in kqueue, or a union of `int/uin32_t/uint65_/void *` in epoll) with the event that you're registering. So when you're notified of the event, you can access this data. In my case, it's a big Conn struct. It contains things like the # of requests on this connection (to enforce a configured max request per connection), a timestamp where it should timeout if there's no activity. The Conn is part of an intrusive linked list, so it has a next: *Conn and prev: *Conn. But, what you're probably most curious about, is that it has a Request.State. This has a static buffer ([]u8) that can grow as needed to hold all the received data up until that point (or if we're writing the data, then the buffered data that we have to write). It's important to have a max # of connections and a max request size so you can enforce an upper limit on the maximum memory the library might use. It acts as a state machine to track up to what point it's parsed the request. (since you don't want to have to re-parse the entire request as more bytes trickle in).

It's all half-baked. I can do receiving/sending asynchronously, but the application handler is called synchronously, and if that, for example, calls PG, that's probably also synchronous (since there's no async PG library in Zig). Makes me feel that any modern language needs a cohensive (as in standard library, or de facto standard) concurrency story.

[1] https://github.com/karlseguin/http.zig*

latch··on Sieve is simpler than LRU
I feel like a simple improvement to LRU is to not promote on every fetch. This adds some frequency bias into the mix and reduces the need to synchronize between threads. Is this a known/used technique?
latch··on Zig cookbook: collection of simple Zig programs that demonstrate good practices
The first example, reading lines in a file, is really slow. Depending on the input, it could be 4x-10x slower. In Go, the equivalent `ReadLine` is implemented on `bufio.Reader`. Thus the implementation can take advantage of the buffering.

In Zig, while this example does use a BufferedReader, the `streamUntilDelimiter` is implemented on the more generic `std.io.Reader` and thus cannot take full advantage of the buffering. It checks 1 byte at a time, and cannot leverage the SIMD-optimized `std.mem.indexOfScalar`.

Of course, you're free to create your own `readLines` that works against a BufferedReader. I do hope we'll see a whole version dedicated to review/polishing of the standard library. And maybe more work on composition, abstraction and interfaces.

The issue is described here: https://github.com/ziglang/zig/issues/17985

latch··on Website search hurts my feelings
If it's supposed to be sorted by name, why does "Ariel Oxy Bleach..." show up before "2-Fold Umbrella..." but "Ariel Antibac Jumbo" shows up after.

This isn't just bad search, it's objectively wrong (and a lot more basic).

latch··on Website search hurts my feelings
The complaint about "rice" is that the tagging is incomplete. "vegan" isn't a bad facet, but the count is wrong. When you select "vegan", most vegan products are removed. I don't see how filtering wrong data isn't a valid complaint.

The problem is obviously that the data isn't tagged. But Costco is huge, profitable and doesn't have a huge inventory, so why can't they tag it? And if the tags are so incomplete, maybe it's better not to include them?

← PreviousPage 2 of 34Next →