HNHacker News
TopNewBestAskShowJobs

pebal

23 karma · joined May 27, 2022

submissionscomments
pebal··on Port of the TypeScript compiler, checker and lsp to Rust, by LLM
Rust could just as well have an optional tracing GC for shared ownership instead of ARC, potentially improving performance rather than reducing it. Having a GC doesn't inherently make a language slower.
pebal··on Port of the TypeScript compiler, checker and lsp to Rust, by LLM
I’m not saying it doesn't work. I’m saying it’s a waste of resources and won’t gain widespread popularity. The current trend is porting source code to languages that generate native code.
pebal··on Port of the TypeScript compiler, checker and lsp to Rust, by LLM
It depends on the workload. ARC can be more expensive than tracing GC, especially with heavily shared objects.

But my main point was that the presence of a GC has nothing to do with how close a language is to the metal. Memory management strategy and low-level capabilities are two separate things.

pebal··on Port of the TypeScript compiler, checker and lsp to Rust, by LLM
WASM makes no sense and has no future.
pebal··on Rust Port of TypeScript (Tsc)
An absurd amount. The cost wouldn't exceed $500 under a subscription model.
pebal··on Port of the TypeScript compiler, checker and lsp to Rust, by LLM
There is no connection. Rust has ARC which performs worse than the tracing GC.
pebal··on A 40ms Go garbage collector pause caused by swap
What exactly don't you understand?
pebal··on A 40ms Go garbage collector pause caused by swap
Manual management involves an explicit function call. Functions called by a destructor are considered automatic.
pebal··on A 40ms Go garbage collector pause caused by swap
Yes, I can, e.g. by implementing deferred deallocation, which is a primitive GC form.

> actually, you don't know

I know because I’ve verified it.

pebal··on A 40ms Go garbage collector pause caused by swap
You might lose 100 ns or 300 ms. When? Nobody knows. You lose nothing when going out of scope with GC.
pebal··on A 40ms Go garbage collector pause caused by swap
It is also unknown how much work will be performed if ARC goes out of scope.
pebal··on A 40ms Go garbage collector pause caused by swap
Because some people are allergic to the word GC and pretend that a duck isn't a bird.
pebal··on A 40ms Go garbage collector pause caused by swap
That "good collector" has higher memory usage, has to pause mutator threads, and makes application code run slower. Thanks, but no thanks.
pebal··on A 40ms Go garbage collector pause caused by swap
It is better to scan everything than to copy it.
pebal··on A 40ms Go garbage collector pause caused by swap
You've described where deallocation can happen, not when it will. Every block exit is a candidate, but which one drops the last reference depends on runtime state: how many other owners the object has and which of them goes away first. With shared ownership across threads it's a race by construction. The free runs on whichever thread happens to release last.

The cost isn't bounded either. Dropping the last reference to the head of a list or the root of a tree frees the whole structure at that block exit, and its size is a runtime property.

And it isn't only block exits. Every assignment to a variable or field holding a reference decrements the old target, and so does removing an element from a container. Swift's ARC doesn't even promise the scope boundary: the optimizer may release right after the last use, which is why withExtendedLifetime exists.

By the same "where" criterion a non-concurrent tracing GC is predictable too, because it can only run at allocation points. That doesn't tell you which allocation will trigger it, just as knowing the block exits doesn't tell you which one will free.

pebal··on A 40ms Go garbage collector pause caused by swap
> The problem with GC in the strict sense is that you cannot predict when it will happen.

It’s the same with ARC. You also don’t know when the counter will reach zero.

pebal··on Garbage Collection: Generational? Incremental? Both -Python Language Summit 2026
The inefficient combination of reference counting and a tracing garbage collector found in Python should be discarded in favor of a fully pause-free, concurrent GC.
pebal··on High-performance garbage collection for C++
The standard provided no GC support. What was there was completely useless.
pebal··on High-performance garbage collection for C++
On your third point (compilers and Boehm): agreed on both counts. The C++11 "garbage collection support" (declare_reachable and friends) never got an implementation and C++23 removed it, so a collector for C++ has to live without the compiler - and Boehm, being conservative on the heap and stop-the-world, is what most people think that has to mean.

It doesn't. I've been building one as a plain header-only library (SGCL, https://github.com/pebal/sgcl - my project): the heap is traced precisely through per-type pointer maps the collector builds at runtime by elimination (a word ever seen holding a non-heap value is data, for good), marking and sweeping run concurrently with the mutators and in parallel over helper threads, cycles are generational, and nothing is ever moved. What it can't do without the compiler is the stacks, which it scans conservatively, like Go before 1.4.

On the "abstract model" point: you're right that it isn't 100% faithful either. The collector reads words of objects while other threads write them, relies on word-sized stores being what every real platform makes them, and reads stacks it does not own. The rules the program has to keep are few (a tracked pointer lives on a stack or in a managed object, never shares storage with data, destructors don't touch peers) and debug builds check them, but it is engineering on top of the platforms, not the standard.

pebal··on High-performance garbage collection for C++
I've spent the last few years on a library-only collector for C++ (SGCL, https://github.com/pebal/sgcl), also non-moving, so a few data points on the "can it compete" question. Disclosure up front: my project, and measured so far only on Apple Silicon/macOS, one machine, nothing tuned.

On the sweep: it's true that a non-moving collector's sweep is proportional to the number of dead objects, and a moving one's isn't. In practice that cost is small and parallel: the sweep walks per-page state bitmaps (about 1 ns per object), runs the destructors of the dead objects, and is spread over helper threads while the mutators keep running - nobody waits for it. What moving actually buys you is bump allocation and locality. Per-type pages with thread-local free bitmaps get allocation to about 5 ns per object on one thread and 9 ns on 24 threads, without a pause and without moving anything; ZGC does 3.6 ns and 26 ns on the same machine. On binary-trees at depth 21, ZGC is ahead on one thread (2.5 s vs 4.5 s) but with a 1.1 GB heap against 340 MB, and on four threads they tie (1.6 s each). Go, which is also non-moving, sits at 6.3 s / 219 MB there. So "non-moving" is not what decides it; the cost of the barrier, the marking and the sweep spread over cores decides it.

On destructors: they run on the collector's threads, in parallel, not on one thread, and the rule is the same as Oilpan's - a destructor must not touch other managed objects, because they may be dying in the same sweep. I don't have a Clang plugin to enforce it statically; there is a runtime check in debug builds and an explicit escape hatch (if_alive()) for the one legitimate case, a destructor asking whether a peer is still there. Oilpan's static verification is the nicer answer to that particular problem.

Where the approaches really differ is how the collector finds the pointers. Oilpan needs a Trace() method per class (or, as someone suggested above, reflection to generate them). SGCL builds a pointer map per type at runtime by elimination - a word that is ever found holding a value that isn't a managed address is data and leaves the map for good - so plain structs with tracked pointers in them just work, at the price of a couple of rules (no union of a pointer with data, stacks scanned conservatively). Different trade-off, not obviously worse.

pebal··on Giving C a superpower: custom header file (safe_c.h)
RC is a GC method and the least efficient one.
pebal··on Should I choose Ada, SPARK, or Rust over C/C++? (2024)
The C++ standard has never included a garbage collector. It only provided mechanisms intended to facilitate the implementation of a GC, but they were useless.
pebal··on There is no memory safety without thread safety
This isn't fully concurrent GC. It pauses mutators threads and delegates them to perform some of the work for the GC.
pebal··on Garbage Collection for Systems Programmers
> I haven't seen a C++ programmer carefully opt into GC for a subset of their allocations even though there are GC libraries written for the language.

Can you give an example of such GC libraries?

> Whoever made that claim? Gamedevs particulary have been writing custom memory allocators since decades precisely because they know free() is not free and malloc() isn't fast.

Game developers use engines based on the GC.

> It's not an illusion, you literally use control over memory management.

The shared_ptr does not provide full control.

> What they don't want is random stalls in odd frames.

You can have fully concurrent GC, without any stalls.

pebal··on Redesigned Swift.org is now live
But that's why Swift generates slower code. Memory is cheap, also for Apple, although Apple would like to hide that fact.
pebal··on Four Years of Jai (2024)
It doesn't matter at all. C4 uses STW.
pebal··on Four Years of Jai (2024)
Azul C4 is not a pauseless GC. In the documentation it says "C4 uses a 4-stage concurrent execution mechanism that eliminates almost all stop-the-world pauses."
pebal··on Four Years of Jai (2024)
There are no production implementations of GC algorithms that don't stop the world at all. I know this because I have some expertise in GC algorithms.
pebal··on Four Years of Jai (2024)
There are none, at least not production grade.
pebal··on Go Optimization Guide
Yes, SGCL is my project.

You can't write concurrent code without atomic operations — you need them to ensure memory consistency, and concurrent GCs for Java also rely on them. However, atomic loads and stores are cheap, especially on x86. What’s expensive are atomic counters and CAS operations — and SGCL uses those only occasionally.

Java’s GCs do use state-of-the-art technology, but it's technology specifically optimized for moving collectors. SGCL is optimized for non-moving GC, and some operations can be implemented in ways that are simply not applicable to Java’s approach.

I’ve never tried modeling SGCL's algorithms in TLA+.

Page 1 of 4Next →