Why is D's garbage collection slower than Go's?
forum.dlang.org
forum.dlang.org
Barriers are pretty cheap; [0] claims a 0.9% time overhead for a card-marking barrier common in generational collection and 1.6% for an object-logging barrier which is also useful for concurrent collection. Apparently they're cheaper nowadays, but the results aren't published yet [1]. That's not to say that the barriers are free, but it seems feasible that collector optimisations could still cover that ground.
It can help throughput, still, by running the collector concurrently. I've felt that while working on a new parallel (but not yet concurrent) collector for SBCL; the program parallelises well, but a serial collector hurts worse than it should by stopping the program.
[0] "Barriers reconsidered, friendlier still!" https://users.cecs.anu.edu.au/~steveb/pubs/papers/barrier-is... [1] https://twitter.com/stevemblackburn/status/14942409060061102...
moonchild's suggestion to monomorphise against GC/explicitly managed pointers might reduce that figure lower; evidently you know D programs better than I, but it doesn't seem unreasonable that the 1% can be won by a faster GC.
Of course, it depends on your coding style, but yes, that is generally true.
You won't notice it on your desktop. But you will notice it in cases I mentioned. Top shelf games also are not going to give up that 1%, or so the game devs tell me.
Unity is also the official developer SDK for Microsoft's HoloLens Virtual Reality Toolkit.
Doing the game in C++ certainly hasn't helped Cyberpunk 2077 performance and being bug free.
- Unity games
- NWN2 which I mentioned
- CrossCode which is an HTML5 game (while being an absolutely great game, performance wise this one was by far the worst I ever experienced, it's laggy on a fucking GTX1070 / 8750H / 32GB RAM system for a 2D pixel-art rpg and made me quit the game in rage more than once)
I don't think this is a reasonable remark.
The whole point is that the 1% penalty is added on top of whatever choices "the likes of Sony and Nintendo" make. If they have alternatives that ensure them a performance win just by chosing the right tool to work with Unity, they won't choose the wrong one.
D had a chance in the games industry with Remedy and lost it, instead a language that supposedly lesser one, is being adopted by major platform owners.
No, it was not a reasonable remark. Pointing someone else's choice changes nothing. The point that you keep missing is that performance is a key decision factor, one among many, and baseline performance penalties imposed by a particular choice of programing is a factor that adds up to all other choices. If you degrade the performance of your offering without any relevant tradeoff, you're making it harder to justify it's adoption.
Feel free to insist in your personal assertion. Those who care about performance feel strongly about gratuitously pile performance penalties without any meaningful tradeoff to show.
Minecraft wouldn't never have happened if Notch was busy discussing if it would be acceptable at all doing it in Java.
Just like too many in HN dream of being the next FANG, too many worry about the ultimate performance when their games would hardly win a fraction of Minecraft when placed into the market.
Getting acceptable performance in my game despite my focus on playability and time to market is one of the benefits of platform choice. I want to focus on my game, not overcoming platform limitations.
I am not saying a single 1% makes a difference but the idea that performance in the platform does not matter is wrong.
Especially that research OSs were written in managed languages, often with better performance!
But nobody said this. Maybe we're too deep in the thread for anyone to remember, so let's refresh our memories. The claim was not that performance doesn't matter, or that 1% never matters, but that any language 1% slower than D is currently could categorically not be considered a systems programming language.
That's even more absolutist and absurd than "performance does not matter".
Uh you know who’s so patiently answering your questions, right? My guess is that he’s been world famous at this since you were a twinkle in your daddy’s eye. He’s had a minute to think about it...
True! Duly upvoted. Although I feel his arguments are closely reasoned and don’t agree with your point. And the person I responded to brought nothing you to the table.
Where is it that Walter is obviously wrong, by the way? Not trying to be argumentative. I just did not see the holes in his argument that you do.
These API have been removed in ISO C++20, only because the biggest C++ GC customers, like Epic and Microsoft, never made use of them and carried on using their own implementions.
Your specific claim at the top of this chain is in response to someone asking "why don't you do this thing [that has 1% performance impact]", to which you respond "I cannot bundle a 1% performance hit in my language, because this would not make it a systems language". Putting aside the usual debate of what a systems language even is, if we consider the ones that are typically not up for debate: C, C++, Rust, these differ in performance by far more than 1% on typical workloads. As some commenters have mentioned, compiler optimizations alone cause pessimizations of greater than 1%, so looking at a number like this as any qualifier of how feasible something is doesn't make sense.
Taking a step back, I feel like you are missing what people actually mean when they're talking about "1%". Like, yes, 1% of Facebook's server load is $$$. Making a dozen 1% improvements to SQLite is a good improvement. But, like, you're conflating this with what you're doing, and it's not at all related. There are companies using Python in production right now trying to save 1%! The reason why this is a "rational" decision is that performance work is not actually a function of how much percent you can shave off your workload, but how you can balance a couple of people trying to wring a couple percent off of your existing code. The unfortunate truth is pretty much all code, even the stuff running billions of CPU hours in datacenters, is leaving tens of percent on the table at the very least just by using a high level language, with poor data access patterns, etc. The reason this is OK is that rewriting all of the code into perfect assembly or whatever is not a feasible task. It would be super tedious, error prone, and require huge amounts of effort. So it becomes relevant for specific people to try shave off a few percent here and there around an inherently inefficient codebase, because that's where the balance lies. Compared to the effort it took to write or migrate the code, the 1% win is always going to be a small fraction of the engineering cost. Otherwise, you'd just replace the thing altogether.
So, circling back around to your point: a 1% loss isn't actually catastrophic. If you put 1% extra code in for no reason and it was easy to get rid of it with something a little smarter, people would rightfully be up in arms. But if you bring actual benefit that is very hard to get any other way, then it's usually going to be welcomed. I mean, people are putting in specialized hardware to slow down their general C++ application code "just" 5-10% to get a fraction of the benefits of actually having memory safety. There's a limit to how much people will tolerate, and it's definitely not like 30%, but if the only penalty was 1% at runtime and you could never have to worry about freeing memory again this would probably be a good tradeoff.
(To answer your bit about why the C standard doesn't have garbage collection: multiple reasons. One is that people who use C have very picky views of how garbage collection should look like on their platform. And the other bit is that garbage collection is typically not "just" 1% CPU overhead, it has interactions with things like latency and memory usage, which I think other people have actually pointed out about this 1% figure. The reason I specifically called you out was because you accepted this and decided to argue for that number, rather than saying "oh, well, there's actually more to it than just that 1%".)
EDIT: you said tons of people. Well... thousands of paying customers in a market much smaller than today's, but that's probably not tons.
SHR R_pointer, 9
AND R_pointer, <cards>
MOV BYTE [R_cards + R_pointer], 0
where <cards> is one below an immediate power of two number of cards, each card covering 2^9 = 512 bytes of heap. This is roughly the write barrier in SBCL on x86-64; there isn't much of a reason to think it will cause a latency spike.On the other hand, "Barriers revisited" does mention a pathology about how cards can introduce false sharing, but that's still more of a smear than a spike.
My guess is that it is not that big of a niche, although the players are probably big
It's like weight on an airplane. Boeing told me back in 1980 that saving one pound was worth a quarter million dollars of investment. That'd be like a million bucks today.
The boeing example is illuminating. One could conceive of infinite possibilities where it would be better to spend the extra million dollars than to get something that saved a pound but didn't meet other project requirements. What good would it be to buy a $1M cheaper 747 if the landing gear broke on each landing? Everything is a tradeoff, even things that are extremely important like weight on airplanes or runtime performance in computers.
Jane Street has used OCaml for about 20 years.
Last year, OCaml merged an improvement to its GC that reduced latency by 75% (by one measure), and reduced execution time by 23% when compiling OCaml and around 6 to 7% when compiling Coq libraries:
https://github.com/ocaml/ocaml/pull/10195#issuecomment-89615...
Performance is definitely left on the table even in mature and mission-critical software.
Remember the stories about traders trying to shave milliseconds off of their trading speed?
I even met a hedge fund that ran their back test platform mostly on hardware because it allowed them to run more tests, which was a strategic advantage.
That said, 1% for free makes the ghost of Admiral Hopper happy and that’s reason enough for all of us.
It's a different tradeoff if the languages primary allocation scheme is GC.
Unfortunately, it's not fast enough.
1% compared to what? Hand-written assembly? C? C++?
Clang and GCC differ by way more than 1% in most benchmarks. Does that mean that one of them can't compile system languages or something?
I mean, I would agree that there definitely is a performance threshold below which you wouldn't consider a language a "systems language" - i.e. it would be silly to write an OS in the language. But 1% seems at least an order of magnitude off.
Sub-optimal performance decisions made by the user are their own business (and, perhaps, their customers'), and not really the language authors' concern. The converse is not true. Sub-optimal performance decisions made by the language authors end up affecting everybody.
I agree that saying a 1% difference would disqualify it as a systems language is hyperbolic. I disagree with the assertion that it's not a big deal. Having a core development team that considers 1% to be a big deal is a very desirable trait in a systems language. Achieving very high performance goals often comes through an accretion of many such 1% (and sub-1%) decisions.
Would Linus accept a language that is more than 1% slower than optimal? Yes, he already does. It's called C. C is slower than optimal assembly.
He chose C because it is easier to write large systems in a higher level language, and it is far easier to write cross platform code in.
There is no control over the duration, or what third party code is doing with them.
But writing at that level of specificity for an entire system greatly increases creation and maintenance time costs. You may lose the opportunity to make future performance improvements as a result.
And they evidently spend very little time on the things that would actually improve the lives of their users in macroscopic ways. The fact that our dominant model of an operating system is essentially c + shell is a tragedy. The operating system does almost nothing for us to help write fast, correct programs in a reasonable amount of time. It is of course for this reason that unix itself has barely improved in a material way in 30 years except in these kinds of microbenchmarks. Third party tooling and languages have of course improved, but largely to cover the gaps that the operating system has largely failed to address in a meaningful way.
Um, when the PC computers switched from real mode to protected mode operating systems, that was an enormous boost to programming. It's hard to understate it.
1% is a lot of leverage. One reason I got paid well when I worked in industry is because I could write code that ran faster. Shaving 1% off execution speed is a big deal.
I don't know but this seems like a trivial decision given how many metrics they must track about the relationship between weight and capacity/fuel cost - either you pay less than you save, or you don't pay it.
> How much would Ferrari pay to make their F1 cars 1% faster?
Probably a lot, but this is a (somewhat literal) Red Queen's race.
> How much would car companies pay to be able to make 1% more fuel efficient cars?
Probably nothing, their precision is already several times that. Empirically, they also don't care much.
> How much would your electric bill save if rates were reduced 1%?
Less than 5 euros a year - though this year maybe more, could be as high as 10. So, negligible - not even worth the time to write this comment, probably.
Overall it seems like "1%" is often a fairly meaningless measure and I'm not sure what any of these examples are supposed to illustrate to me about compiler design or language choice since the context for each seems to matter more than anything else.
Then apply this 1% saved logic to everything else, food, clothing, leisure, insurance, medical bills, sport
1% is a lot, cumulatively
If you want to make a point about where 1% efficiency matters, go ahead and make it! I might even agree! But it still won't have anything to do with systems programming languages.
your system being 1% faster mean everyone who depend on your system will be 1% faster, and if you yourself apply that logic to your system, it's commutative, and the results ends up being impressive
if you don't understand that, then there is no point arguing further
What the fuck?
A quick look at the paper indicates that researchers were able to find cases in which the cost was as low as 0.9%. I don't think it's a good idea to operate as if that is the cost.
0. https://forum.dlang.org/post/mpczwoeykpcwjakuderd@forum.dlan...
1. https://forum.dlang.org/post/eujlhvszhpbyoxemscxb@forum.dlan...
2. https://forum.dlang.org/post/ssq34d$2ur7$1@digitalmars.com
Besides, I seriously doubt that there's much gold in unambiguously recognizing stack pointers. I try to put as much as possible on the stack, but it just doesn't seem like much. The more stack allocation you manage to do, the more code bloat gets generated. Any malloc'd pointer will look just like a GC pointer, so you'll be paying the penalty for all of those.
How these tradeoffs play out will be difficult to determine in advance. If you've the confidence that it will work out favorably, by all means, give it a try.
Of course, one could add user annotations, but annotations are generally disliked.
I don't see why. A small amount of manual tagging for extern(C) things—done in this case as part of the standard library, so opaque to user code—will go a long way in that regard. And in any custom allocator, automatic inference should work just fine, since the data returned will be clearly derived from the already-tagged malloc/mmap/VirtuallAlloc/whatever (type info does still have to be associated manually, of course, as part of the GC.addRange dance, to avoid spurious pinning).
You'd also have to keep around a massive amount of data at compile time. A variable that's a pointer to a struct, can't just be represented as a pointer to a struct. It'd have to have a map of which pointers in the struct are managed for every pointer to that struct.
The way the D compiler is implemented now, a struct S has only one instance, as every pointer to a struct's type points to that same instance. And still sometimes people compile programs so large the compiler runs out of memory.
You should revisit that corner of the thread to revive your memory. The discussion was about adding an extra 1% performance penalty for reasons, not the importance of a 1% delta either way.
You could always use huge pointers in the general case, and treat 16-bit pointers as an optimization only. This is probably how these platforms might be best supported by something like modern GCC/LLVM.
Data flow analysis isn't that good.
It definitely wasn't in the old DOS days. But it's worth testing the approach in the context of modern platforms.
Go's GC is non-moving and it can't safely be made moving in the current state[0].
It also used to be partially conservative until 1.3[1][2][3]: the heap was precise, but the stack was conservative.
[0] https://github.com/golang/go/issues/46787
[1] https://go.dev/doc/go1.3#garbage_collector
[2] https://docs.google.com/document/d/1lyPIbmsYbXnpNj57a261hgOY...
[3] https://docs.google.com/document/d/13v_u3UrN2pgUtPnH4y-qfmlX...
[0] Fast conservative garbage collection https://users.cecs.anu.edu.au/~steveb/pubs/papers/consrc-oop...
I was astonished to learn that researchers found a way to implement a compacting malloc (!!!) by using very clever virtual memory tricks - and which they were able to use to demonstrate memory usage improvements in a long-running Redis instance that used their drop-in malloc replacement:
Optional GC, also means you are free to go for it manually, if there is really and truly some bit of extra performance needed. In regards to language wars, it's arguably better to make the case of how to make manual memory management more convenient to use, when the optional GC is turned off.
“Convenience” makes it sound like syntax sugar, it’s not merely convenient.
It eliminates certain types of bugs fundamentally, often it can simplify how the code is reasoned about.
Just like static typing is a convenience. And dynamic typing is a (different) convenience. And unit tests are a convenience.
The computer runs fine without any of these. And none of them are a magic wand that eliminates all bugs. Instead all of them make debugging easier by reducing the search space, usually eliminating fairly trivial bugs. Which is more useful than it sounds because a lot (the majority?) of bugs are pretty trivial and at the same time hard to find because they are often so trivial that they stare us in the face and we can't see them.
But again, all these things are conveniences for the human programmer, they don't make one iota of a difference for the actual computer running the bits.
Even if language X loses in microbencharks against language Y, if the application written in X is still within the project delivery SLA for acceptance testing, who cares about the microbencharks.
But maybe the view of modern society as a tower built on the back of a million million one time conveniences that have since become indispensible isn't that far from the truth either?
As a real world example, consider a firmware boot loader I wrote for a previous client. They needed to apply code polymorphism (think, extreme ASLR) when they applied an over-the-air update or a factory reset of firmware. This is part of their defense in depth strategy: if a gadget attack were discovered in their system, they wanted to mitigate its effectiveness across all deployments. The problem with this approach is that boot loaders are tightly space constrained. To implement a sound code polymorphism strategy, around a dozen different heuristics must be applied in several passes over the code. These heuristics had to be weighed against timing and cache coherency constraints in the code under mutation, but most critically, these heuristics had to fit in the very limited space set aside for the boot loader. The GC was an optimization. It significantly reduced the code size of the heuristics.
The last 50 years of CVEs suggests they're not so trivial.
That's why static typing and unit testing are both practically much better at reducing bugs than they have any right to be theoretically.
Java is one of the languages to blame for such misunderstanding, many other languages, even Lisp variants like Interlisp and Common Lisp, provide all the tooling to manage resources like one would do in C and co languages.
Having a GC doesn't preclude those other features.
On D's case, we have the problem that language designers are against moving the GC forward, as they refuse to introduce managed pointers alongside untraced pointers.
So even optimizations that Mesa/Cedar, Modula-3 and Active Oberon were capable of before D came to be, aren't possible in D with its current language design.
This is incorrect. There even was a concurrent collector written for D, but it failed for technical reasons, not anyone blocking it.
The idea that I would block any improved GC implementation for D is just silly. D is free and open source and Boost licensed. I couldn't possibly stop anyone from doing a better one.
You're free to prove me wrong about what can be done, I would welcome your implementation!
C# already offers what I care about in managed languages with C and C++ like capabilities.
And honestly why would anyone go through the trouble to prove you wrong, if in the end it eventually has the same fate as the concurrent one, that only works in Linux anyway.
I'm curious why you don't use Managed C++ (the one with separate GC pointers)?
Managed C++ was replaced by C++/CLI in .NET 2.0.
It is the best way to write bindings to native libraries on the .NET ecosystem without having to bother with getting P/Invoke or COM bindings attributes quite right.
However it is constrained to Windows, so for portable .NET code, P/Invoke it is.
In fact C# actually has three models of pointers, regular references, JIT intrisics like IntPtr/UIntPtr and raw unsafe pointers.
p.s. those seem to be persistent object focused. Persistence in a caching memory model would be a side-effect (and doesn't even have to be used). Do you know if there exists research primarily focused on 'a caching memory model' as an alternative to (general) GC or reference counting strategies?
The general idea is pretty straight forward, you have a ~fixed size chunk of (virtual) memory that active memory objects reside in. Garbage and very rarely used references get flushed to disk. Presumably, if the operating memory requirements (i.e. active objects) are met by available Ln cache layer, this should be a viable alternative to 'collecting garbage'. We're trading the overhead of tracing/ref-counting with the cost of cache mechanism. If cache is tiered, then the 'garbage' & 'rarely used' objects will end up in nice VM blocks that OS will flush to disk. One approach basically requires the use of one annotation at language level to distinguish 'long lived' objects.
In the beginning the biggest drawback seemed to be performance, so people have been hell-bent on fixing that for decades now. Completely ignoring that the "don't need to care about memory" mindset is flawed. You always need to think about memory, and if you do you might as well do it yourself. It doesn't take more time to do in the long run and you are forced to design a better system because of it.
For scripts it is a really nice have. But for a proper language writing large-scale applications? No thanks.
GC will always be needed to solve problems dealing with general graphs containing possible cycles. Not coincidentally, this describes many of the GOFAI problems for which GC (in the context of LISP) was first developed.
Weak and unowned pointers are pretty simple to deal with, and I’ve never seen a situation where it’s really confusing whether a pointer should be strong or weak in swift. Everything mostly just works, and if you need to point back to the object that owns you, it’s usually pretty clear by the abstraction you’re building that you need a weak/unowned reference to do it.
That's not the interesting case; that is indeed easily addressed by "weak" pointers. What's interesting is the case of a spaghetti reference graph where none of the objects definitely "owns" any other, and even the choice of GC roots might be dynamic.
With Rust and C++ (don't know about Swift), you still have to spend a lot of time thinking about memory.
You don't think about memory, you just think about how to ensure that you're always referencing live objects. And you need that kind of thinking anyway to deal with non-memory resources, where a GC won't help you at all.
There are other ways such as customozong new/delete.
Depends on your situation what you can or cannot do/afford.
Modern C++ unfortunely still has too much people writing classical C with classes, to make it properly safe.
Rust, yes, affine types are great, but their usability is only justified in scenarios where having any kind of automatic memory management isn't acceptable.
Well, it would be a bad argument :D
A reference counter is arguably the slowest option, so Swift is out of the picture (besides it leaks circular dependencies). Rust is a good idea, but it is a tradeoff. Quite a lot of programs can’t be expressed with strictly tree-like lifecycles, and then you are either left with the yet again slower (A)RC, or unsafe. C++ is in pretty similar shoes. They are good for the niche they were made, systems/low-level programming.
Let me quote from the Garbage collector handbook: “above all, memory management is a software engineering issue. Well-designed programs are built from components [..] that are highly cohesive and loosely coupled. [..] modules should not have to know the rules of the memory management game played by other modules. [..] GC uncouples the problem of memory management from interfaces.”
It says nothing about not thinking about memory, you are free to optimize it where you need it.
EDIT: Oh and dynamic allocation is much faster with a modern GC than with malloc and friends, plus a moving GC can defragment heap.
It's not that simple. Obligate refcounting has lower throughput than obligate tracing GC but offsets that with more predictable latency, which might be good for the Swift use case. Rust uses a mix of stack-based allocation, manual heap allocation and RC - and the RC only needs to manage "owners" as determined via the borrowck, as opposed to ephemeral references. So is way more efficient than something like Swift ARC, where refcounting is essentially ubiquitous and any operations via references will incur slow atomic RC updates.
That’s shared pointers recursively freeing up their reference trees. There is a great paper showing that RC tracks object death, while tracing GCs track object “liveness”, so its essentially two sides of the same coin.
As for RC, sure, at compile time many of the work can be omitted, but atomic writes needed for the counters are fundamentally slow operations on modern architectures.
Haskell needs a GC too. You can't do ergonomic functional programming without GC (it has been attempted).
The problem however is, actually 2 problems:
- stop the world, nobody wants that in a world with lot of cores and threads
- it doesn't scale, the more pointers in your heap, it'll need to scan and traverse your WHOLE heap whenever it needs its buffer to grow, and that doesn't scale well
So it's good when you don't have much in your heap, and it starts to loose its benefits the bigger your program become, i wouldn't use it for my servers
But that's not the main problem of D, since the GC is optional, it's just not competitive with what's available in the market today
The people who want to drag D into the Java/C# territory are the problem in my opinion
D would be better if it focused being a system language, and took what C had to offer and put it to the next level, simplify the language, boost the existing features, allocators, pattern matching, tooling, compiler performance, hot-reload, binary patching
That's the thing i want to hear when there is a new version, not the endless GC topics
Absolutely and that's the majority of D community.
> took what C had to offer and put it to the next level
ImportC is a fantastic stuff that Walter is working on.
> simplify the language
Sane defaults, but the ship has sailed. Reminds me of a talk Scott Meyers gave a while ago "The last thing D needs (is to hire him)". I think it's time they hire him.
> compiler performance
Rather focus on just LDC or GDC and drop DMD altogether. DMD is a good piece of software, but for such a small community, I find it alarming that they waste human effort across 3 different compilers and still complain lack of resources.
I agree, it's one of the things that stands out when you decide to pick a system language: "how does it play with C? can i easily consume the ecosystem?"
> Rather focus on just LDC or GDC and drop DMD altogether. DMD is a good piece of software, but for such a small community, I find it alarming that they waste human effort across 3 different compilers and still complain lack of resources.
I disagree, there is value in having your own backend, DMD compiles so fast, it's a comparative advantage, they should never give that up
GDC/LDC are great because that allows D to be highly portable, even if they are slower to compile than DMD
Even Zig people decided to maintain their own self hosted backend for that reason, performance and independence
They learnt from D, a real language has its own backend, if you don't then you are just LLVM sugar
So you'd use DMD for development because it compiles fast, and once debugging/testing/etc are completed, you build the "production release" using LDC/GDC?
No problem with LLVM, but multiple implementations is healthy. And in any case I'd probably think it makes sense DMD came into existence alongside the idea of D first seeing Walter's experience (and well lack of LLVM back then).
D doesn't even have hardware vendors shipping commercial embedded SDKs for baremetal development like Java, C# and Go have in 2022.
I don't entirely agree, I think having a garbage collected pointer (GCP) can make a lot of sense as a performance optimisation: manual memory management (MMM) and reference counting (RC) have stampeding characteristics similar to GC pauses when releasing large hierarchies e.g. persistent data structures, or large trees of widgets: all the tree is freed synchronously and recursively, if it's large it can be very sensible.
This stampeding is more predictable than a GC pause, but it's no less problematic when it occurs, and mitigating it can be difficult (you have to manually shunt the objects you'd like to release off-thread). A GCP can do that shunting on its own, letting a GC thread perform the releases asynchronously, and piecemeal.
Furthermore, GCP means you're not expecting to precisely track allocations at the system level, which means the usual tricks (arenas, bump allocators, ...) can be applied by default to GC pointers, where they usually can't be to MMM or RC pointers. This decreases allocation and deallocation overhead by reducing the amount of work the system has to do.
You could also use async programming to do the same thing in GC-free code. Or just use an arena that can be freed all at once.
No? Allocations are blocking, "async programming" won't do anything.
Unless you have an async dealloc which does the shunting implicitly, at which point you don't need an async dealloc, you can just have a sync one which shunts the actual dealloc and mandate a background thread. Which means now you're mandating a background thread for freeing memory. And you hope that the load is light enough that implicit thread which you've tasked with all deallocations (rather than just the ones which make sense) keeps up.
> Or just use an arena that can be freed all at once.
That is essentially the same thing, you now have a different allocation strategy for these, except it's a lot more limited and specific.
Modern malloc/free implementations must satisfy requirements such as being scalable. It eliminates any chance to implement blocking deallocation from non-owner threads. A modern implementation already has to be async in some ways. You can check paper behind Mimalloc for example.
Your objects in trees and arrays don't have the same lifecycle?
Quoting GP:
> ... reference counting (RC) have stampeding characteristics similar to GC pauses when releasing large hierarchies e.g. persistent data structures, or large trees of widgets: all the tree is freed synchronously and recursively, if it's large it can be very sensible.
This is mostly where RC & MMM fails. This is also where you are supposed to use arena allocators.
If you have a tree you're constantly adding to / deleting from, you can use a Pool backed by an Arena. At that point you don't use pointers, only indices into the pool. When you want to remove something from the tree, you mark the objects in the pool "deleted" until you want something added back again.
No, there is a nested set of lifetimes.
Take a hierarchy of widgets for instance, if you change a sub-view you're going to reclaim the widgets composing that sub-view, replacing them with widgets of an other sub-view.
Then there's persistent datastructures, where lifetimes are reversed (higher nodes are younger), you update a node, you're going to invalidate all the nodes on the path to that, but the nodes which are not on that path are still alive.
> If you have a tree you're constantly adding to / deleting from, you can use a Pool backed by an Arena. At that point you don't use pointers, only indices into the pool. When you want to remove something from the tree, you mark the objects in the pool "deleted" until you want something added back again.
So you're implementing an ad-hoc GC by hand (and so does everyone else) and making your traversal more complicated.
If there's a GC pointer provided (by the language or a library), in the vast majority of cases it will work well enough and you won't need to waste time on that, you can solve actual issues instead.
Why? The gc is optional in D. You can use malloc/free, or your own custom allocator.
There are other aspects where D gets hit from both sides because it never picks a side, instead it tries to cater to both sides, never going full into any of them. I guess there is some benefit in a language that's general and doesn't force you into specific paradigms, but it also increases the surface area (maintenance area) of the language and adds complexity for developers when every step of the way you have to consider the alternatives.
Walter has a misconception here - Java does scalar replacement - it doesn’t currently do stack allocation.
Do I understand it correctly that you are saying real stack allocation would involve allocating the whole object, including the header, on the stack, and passing references to such objects to (non-inlined) functions that worked on either stack allocated or heap allocated objects?
No, they become data flow edges, so could be in a register, or part of an addressing operation, or value-numbered, or nothing at all if they’re never used, so only on the stack as a worst-case fallback.
> Do I understand it correctly that you are saying real stack allocation would involve allocating the whole object, including the header, on the stack
Yes, which is useful in some cases, but generally a lot weaker than full scalar replacement.
Part of the issue can be that languages that do not provide any options for memory management, can try to make it seem that GC is more of a liability than it is. In the case of Nim, D, and other langauges... They are giving options, versus none. The lack of convenient choices, might be the greater issue, versus stigmatizing GC.
That's why we have different programming languages. The danger of making a language a jack-of-all-trades is that it will be master of none. Restrictions are more often than not a good thing, they give the power to reason, for both humans and machines (giving, for example, memory safety).
When D was originally developed it was intended to be a language that was in between C and Java. D was to be a very low level language. So, it was not as low level as C and not as high level as Java. It still managed to outperform Java and rival C. Ten years from now it'll be even faster.
As anything but a compiler expert, I don't understand how GC is a necessary condition of having memory allocation within CTFE[0], can somebody expand on that?
[0]: CTFE seems to be an acronym for https://en.wikipedia.org/wiki/Compile-time_function_executio... (I was not familiar with it).
2. Zig and C++ require any code to be run at compile time to be specially marked (comptime, constexpr). D will run any code that appears in a const-expression at compile time.
3. Zig and C++ require that if a function is to be used for CTFE, the entire function must be compatible with CTFE. D only requires the path taken through the function to be compatible with CTFE. D even has a `__ctfe` pseudo-variable that can be used to branch within a function to compile-time and run-time paths.
There are proposals to extend it, but there are some non-trivial const correctness issues to be solved there, AFAIK.
fn double(n: usize) usize {
return 2*n;
}
const MyArrType = [double(5)]u8;
Similarly, point 3 is not really true either, what matters in Zig as well is the path that the function took during evaluation. const std = @import("std");
fn double(n: usize) usize {
if (n % 2 == 0) {
std.debug.print("even!", .{});
}
return 2*n;
}
const MyArrayType = [double(5)]u8; // this works
//const Bad = [double(6)]u8; // this will fail to compile
You are right about allocation not working during comptime evaluation at the moment, but this is not a final design decision, just the current status quo.The bigger issue D has with this example is that normal parameters are always runtime. If you wanted this to work in D you would need to implement 2 versions of "double", one that takes n as a template parameter and one that takes it as a runtime parameter. D keeps comptime and runtime parameters separate, making comptime-knowness a "parameter color" which in practice means having to implement things twice if you want it at comptime and often in different ways. There are some things that can work with both but it's small subset of both.
WRONG. if in a CTFE function works without problems. static if has nothing to do with CTFE. static if is conceptually a beefed up proper #ifdef/#endif.
The biggest issue D has is all the FUD and misconceptions that are propagated about it.
C++20 has limited support for allocating at compile time.
I mostly write native mobiles apps (hello Kotlin), but ocasionally deal with backend codes in... Go.
If I need "C++ replacement" for system programming, then the choice would be umm... Rust?
You can call the mmap syscall (or a wrapper like C.malloc) from Go just fine, and obtain non-GC pointers. Not all Go pointers are pointers into the GC-managed region. This makes me question the rest of the opinions herein.
Erlang has this model with it's lightweight processes. And it's a great model that helps not only with GC, but also guaranteeing no shared state between different parts of the code.
That would mean disallowing cyclical references across different GC heaps; non-cyclical references would just create additional GC roots. It would be a step away from totally automated memory management and towards something closer to "smart" reference counting.
Accounting for memory allocation and object initialization (construction etc)?
Legit question coming from a D enjoyer
I've implemented novel GC algorithms in firmware. There are some algorithms that are just more compact when GC is available. Algorithms and heuristics that are graph heavy or require pruning nodes with possible cyclic references are just more elegant and code space efficient with GC. However, firmware is highly resource and timing constrained, which means that the GC is specialized.
Collectors can run on a separate core and show double the speed, but then you have to halve the total number of working processes on your machine. Thus total work done per cycle remains the same.
Or it can just be lazy and run collection half the time, and thus show half processor use. But RAM use will then double for the same work load. At the end actual work done per byte of memory remains the same.
So focus on efficiency. The user can utilize the spare cores and RAMs for paralyzing their work, or for more work.
Basically, you should only use GC for small litter, i.e. some short-lived tiny objects, and for larger and longer living objects use deterministic memory management tools available in the language and libraries.