Major standard library changes in Go 1.20
blog.carlmjohnson.net
blog.carlmjohnson.net
It is unclear when, if ever, we will pick up the idea and try to push it forward into a public API, but it's not going to happen any time soon, and we don't want users to start depending on it: it's a true experiment and may be changed or deleted without warning.
The arena text in the release notes made them seem more official and supported than they really are. We've deleted that mention from the release notes to try to set expectations better. Posting here to leave a note for people who are curious where they went.
Ie., the next frontier is moving from RAM to cache, since CPUs are not going to get faster than programs are "already slow".
If you rewrite some OOP/Pointer-Machine-Model/RAM-Thrashing programs for modern CPUs, you can get 100-1000x speed-up.
The evolution in language design seems, then, to be about exposing more of the hardware to enable developers to target it effectively.
- normalized, small/tight data structures
- data structures that are closely related in computational terms
- smaller interfaces and functions that tend to be more re-usable and general
- fewer if/else branches and more existence based branching via loops
- fewer "business level" generics, macros and similar abstractions, because you can dispatch easily via tagging to concrete types
- less code that "digs"/"drills" into data and more code that composes data
- generally a simpler (less coupled) end result
This all comes with a cost of having to do upfront design and exploration in order to decompose and lay out your data. And while it reduces the mental overhead of understanding the individual pieces of your program on a day by day basis, it might increase learning curve of seeing the big picture, especially at the beginning. So it is a tradeoff.
But I think it would be too dismissive to say that the programmer is doing compiler work here. It is design work, and it is unlearning some of the notions of how to structure programs that carried on since the 90's. Some of which are performance related and some of which are about sensible code structure.
For a very simple comparison, I recently was testing a (poorly) custom built data-oriented Entity-Component-System for usage in games with a more typical "componentized" object approach. No multithreading or anything complicated.
On my system, the typical approach could generate about 1000 new objects and attach a single component in about 1 millisecond.
The data-oriented approach could generate about 100,000 new "objects" and attach a single component in about 0.5 milliseconds.
Same thing in the end, but one is roughly 200x faster in the same time frame. It's pretty stunning when you see stuff like this in benchmarks.
I could be wrong in the assumptions - but OP is talking about fitting stuff in CPU cache don't really see how that translates to your scenario.
-----s.-ms.-us.-ns|----------------------------------------------------------
0.1 ns - NOP
0.3 ns - XOR, ADD, SUB
0.5 ns - CPU L1 dCACHE reference (1st introduced in late 80-ies )
0.9 ns - JMP SHORT
1 ns - speed-of-light
?~~~~~~~~~~~ 1 ns - MUL ( i**2 = MUL i, i )
3~4 ns - CPU L2 CACHE reference (2020/Q1)
5 ns - CPU L1 iCACHE Branch mispredict
7 ns - CPU L2 CACHE reference
10 ns - DIV
19 ns - CPU L3 CACHE reference (2020/Q1 considered slow on 28c Skylake)
71 ns - CPU cross-QPI/NUMA best case on XEON E5-46*
100 ns - MUTEX lock/unlock
100 ns - own DDR MEMORY reference
135 ns - CPU cross-QPI/NUMA best case on XEON E7-*
202 ns - CPU cross-QPI/NUMA worst case on XEON E7-*
325 ns - CPU cross-QPI/NUMA worst case on XEON E5-46*
|Q>~~~~~ 5,000 ns - QPU on-chip QUBO ( quantum annealer minimiser 1 Qop )
10,000 ns - Compress 1K bytes with a Zippy PROCESS
20,000 ns - Send 2K bytes over 1 Gbps NETWORK
250,000 ns - Read 1 MB sequentially from MEMORY
500,000 ns - Round trip within a same DataCenter
?~~~ 2,500,000 ns - Read 10 MB sequentially from MEMORY~~
10,000,000 ns - DISK seek
10,000,000 ns - Read 1 MB sequentially from NETWORK
?~~ 25,000,000 ns - Read 100 MB sequentially from MEMORY~~
30,000,000 ns - Read 1 MB sequentially from a DISK
150,000,000 ns - Send a NETWORK packet CA -> Netherlands
1s: | | |
. | | ns|
. | us|
. ms|
(https://stackoverflow.com/a/33065382)However,
0.001 ns light transfer in Gemmatimonas phototrophica bacteriae
biology has much more performant/optimized machines, therefore, yes, plenty of room for improvement in silico.It's not clear that organic solutions at that level can do programmable computational work, nor that their work is at all deterministic.
At best, it would seem the organic direction for computing will be about building robots rather than CPUs.
The quotation marks around organic are just there to point out that there is something wrong with the dichotomy organic (various pro/eu-karyotes from bacteria to humans)/inorganic (from thermostats to CPUs).
[1] Michael Levin: Anatomical decision-making by cellular collectives https://www.youtube.com/watch?v=Z-9rLlFgcm0
Just on the risks of early miscarriage from wrong number of chromosomes I'd say your numbers are way off.
> Miscarriage is the most common complication of early pregnancy.[21] Among women who know they are pregnant, the miscarriage rate is roughly 10% to 20%, while rates among all fertilisation is around 30% to 50%.
https://en.m.wikipedia.org/wiki/Miscarriage
So 30-50% failure rate.
[1] https://dictionary.cambridge.org/dictionary/english/newborn#...
[1] https://www.anandtech.com/show/13905/tsmc-chip-yields-hit-by...
I'd say a more convincing argument for deterministic machinery is identical twins - I don't know how much variation there is to put in numbers.
Yes, that is the point: biodevelopment is deterministic and computational. Watch the Michael Levin video linked above: they cut the head of a planarian worm, it grows back a head; they cut the tail, it grows back a tail; they cut the tail and the head and change the bioelectric gradients, it grows back two heads or two tails.
https://www.cdc.gov/nchs/fastats/infant-health.htm
Number of infant deaths: 19,582
Deaths per 100,000 live births: 541.9
Leading causes of infant deaths:
– Congenital malformations, deformations and chromosomal abnormalities
– Disorders related to short gestation and low birthweight: not elsewhere classified
– Sudden infant death syndrome
This number is far too high. The rate of conjoined twins (violating "1 head") is about 1 in 50,000 [1], and the rate of "limb reduction defects" (violating "2 hands and 2 legs") is about 1 in 1,900 [2].
Those correspond to 99.998% and 99.94% respectively. 3-4 nines is still impressive for such a complex system, but let's not claim it's 7+ nines.
[1] https://www.chop.edu/conditions-diseases/conjoined-twins [2] https://www.cdc.gov/ncbddd/birthdefects/ul-limbreductiondefe...
[1] Importance of Angiographic Study in Preoperative Planning of Conjoined Twins Case Report, https://www.sciencedirect.com/science/article/pii/S180759322...
Then why not simply give the correct, still impressive, figure, as I suggested?
> the regeneration is always, 100% a head, if no change in the bioelectrical gradients
This is also a meaningless statement. It's correct 100% of the time, except when something goes wrong and it's not.
Can you quantify the likelihood of something going wrong with the "bioelectrical gradient"? I'm not familiar with this organism but I suspect it's several nines, but less than 7.
In general, probabilities less than a certain amount stop being meaningful, because it's more likely that the model used generate the probability fails to reflect reality. See https://www.lesswrong.com/posts/AJ9dX59QXokZb35fk/when-not-t...
The change in the bioelectrical gradient is a human intervention over the organism. Watch the video I linked above. There is 0% chance of "something going wrong with the bioelectrical gradients", it's at the experimenter's will. If you are not familiar then why do you suspect? Your statement is not even meaningless.
Okay, great, we're getting somewhere. So you concede that the true number of human birth defects is on the order of 4-5 nines.
We know this because we've observed a huge sample size of human births. Meanwhile, the experiment you reference only observed a small set of planarian worm amputations. So we can't conclude there are even 4-5 nines of reliability there, let alone "100%".
Otherwise, we could simply observe a few hundred human births, observe no defects, and conclude that human births are also "100%" reliable.
In our original debate, we were both slightly wrong about the number of nines of reliability in human births. However, you are now infinitely wrong by claiming an infinite number of nines of reliability in planarian worm amputations. I don't know whether the actual number of nines is 5, or 10, or 20, but I can be certain that it's not infinity, because that would violate the laws of probability.
Again, you have no idea what you are talking about, as you admitted you are not familiar with the planarian worm organism and regeneration research, and it's not a problem, we are all ignorant about various things, that's why we learn: too bad your learning appetite has been a casualty to the illusion of LessWrong "rationalism". Nevertheless, it is really funny to see you being "rational" and speculating upon things you have no understanding and no desire to learn about. I really laughed reading your now deleted comment starting with "Zero is not a probability."
Just to make it clear for anyone else who might read this: it is impossible to throw a ball in the air and see it flying in the air forever. There is 0% chance of that ever happening. There are no "laws of probability" to be violated in this "experiment". Just the same, when you amputate a planarian worm head, regardless if you did it once, never, or 100,000 times before, it will always 100% regenerate a head, if you, the experimenter, haven't altered the bioelectrical gradients of the worm [1]. The planarian worm regeneration is still being researched and it is revealing biology as a deterministic computation in the morphospace with abilities far exceeding what we currently can muster with our CPUs.
[1] Planarian regeneration as a model of anatomical homeostasis: Recent progress in biophysical and computational approaches, https://www.sciencedirect.com/science/article/abs/pii/S10849...
Legit question, I have no idea about how this would behave.
I know, sounds crazy. Who’s gonna make such a drastic change to the industry?
Expanding the cache so everything fits in it is one way to achieve the performance uplift of cache hits, but cache is expensive compared to memory, and current cache sizes are tiny compared to RAM. If cache gets 50x bigger tomorrow, chips will get at least 10x as hot, power-hungry, and expensive.
Look into Apple’s unified memory.
> ... the available RAM is on the M1 system-on-a-chip (SoC).
[1] my understanding is that RAM and CPU processes are very different and it is hard to produce a chip with both features while remaining optimal.
I want an alternative to Python that can provide the same awesome batteries-included experience. I am also considering Nim but I want a language with a strong industry backing.
But to answer the question: yes, Go has support for SMTP in net/smtp and a lot of different serialization formats in encoding/
You can browse it all here: https://pkg.go.dev/std
> what about support of smtp
Yes, https://pkg.go.dev/net/smtp
> data serialization formats
What you can find in the "encoding/*" subpackages: JSON, XML, Gob, CSV, ...
My team uses Go whereas the rest of the company heavily uses Python. Our vulnerability scanner tool detects hundreds of high score CVEs just in their container images. Comparably there have been times I haven’t updated our distroless base image for a year and there isn’t even a single vulnerability (this one: https://github.com/GoogleContainerTools/distroless/blob/main...)
In terms of defending your software supply chain, eliminating the cruft that is required to run an interpreted language in a container make a a huge difference.
Go has a runtime, of course, but it’s part of the binary.
Actually, the single binary alone is worth the effort. But it goes much deeper than that.
Edit: or will ever be. It is definitely explicitly not an equivalent of go's net/http. Indeed, there is probably never going to be an equivalent of net/http in the Java stdlib (since they prefer to rely on the user choosing one of the existing server frameworks, such as Jetty).
> Provide a command-line tool to start a minimal web server that serves static files only. No CGI or servlet-like functionality is available. This tool will be useful for prototyping, ad-hoc coding, and testing purposes, particularly in educational contexts.
> It is not a goal to provide a feature-rich or commercial-grade server. Far better alternatives exist in the form of server frameworks (e.g., Jetty, Netty, and Grizzly) and production servers (e.g., Apache Tomcat, Apache httpd, and NGINX).
https://docs.oracle.com/javase/8/docs/jre/api/net/httpserver...
I see it's still available after the modularization effort, and it is:
https://docs.oracle.com/en/java/javase/18/docs/api/jdk.https...
That server is not remotely comparable to the one in Go.
The Java one is not usable in anything beyond hello world, and is explicitly not intended to be.
I mean it would be trivial to implement that reverse proxy in Go. And I do mean trivial; Go also includes a reverse proxy utility, so you can implement something basic in about 5 LOC.
At this point it’s hard to believe you’re being genuine.
A weekend project and going at scale isn't the same thing.
https://docs.oracle.com/javase/8/docs/jre/api/net/httpserver...
Compared to sat Rust Go has a very large and practical std lib ! You can actually do something with it I/O wise.
To save some of my karma points, Rust do have a big "community with crates"
My first time with Go was one of the rare experiences, where I just wrote code in a new language (some cryptography, some interactions with rest APIs), and it just worked. No wrestling with obscure features, no hidden magic.
Currently I use it very often for various side projects.
> These are predefined layouts for use in Time.Format and time.Parse. The reference time used in these layouts is the specific time stamp:
01/02 03:04:05PM '06 -0700 (January 2, 15:04:05, 2006, in time zone seven hours west of GMT). That value is recorded as the constant named Layout, listed below. As a Unix time, this is 1136239445. Since MST is GMT-0700, the reference would be printed by the Unix date command as:
Mon Jan 2 15:04:05 MST 2006 It is a regrettable historic error that the date uses the American convention of putting the numerical month before the day.
Using the American convention is regrettable, but putting the year after the time is even more regrettable IMHO. Not sure which timestamp format does that? Plan 9?
I think it's comparable to Python.
Go's stdlib was one of the reasons I ended up with Go instead of Rust (might have changed; Rust had a lot community content, but not a comprehensive stdlib; last checked 3-4 years ago).
Still the case now (depending on your definition of ‘comprehensive’ of course). It’s an explicit non-goal of Rust to include “everything” (eg. http, crypto, random numbers) in std because of the stability promises - you can’t make breaking changes to std unless you’re fixing a soundness issue AFAIR.
A stdlib must not contain everything but a solid cryptography lib is probably a good idea.
I would still like to have a more comprehensive or high level stdlib for Rust that is maintained by a core Rust team.
Idiomatic Go eschews frameworks and - as a former Java developer - that is something I really like about it.
In Go you can write a production ready, well tested, _concurrent_ web application with routing, auth, sql storage, html templating, image optimization, and so on without fetching third party libraries. And you’re not leaving official docs for it.
I mean, you should hear the screams from Rust/Go people about how bad C is, and now you're just going to do it?
"you can’t import it by default", "you probably shouldn’t be using it at all", "experimental arena package", and 'you must opt-in to even be able to use it via an GOEXPERIMENT environment variable'
As for screams from Rust/Go people about how bad C is, it said:
> This is highly efficient, but also highly dangerous. What if the programmer makes a mistake [...] > > To mitigate the risk of these kinds of bugs, the arena package will deliberately cause a panic if can detect someone reusing memory after it has been freed.
and goes on to explain that each arena has its own unique/distinct address space, so that if a pointer to an object still exists and is dereference it will get a memory access fault, causing the Go program to terminate, with an error message specifying that arena as the cause of the problem.
Sounds like a pretty reasonable/cautious experiment to me.
As a rust student myself, even the unsafe{} block is stupid
But the great thing with unsafe blocks is that you get to create a safe API abstraction on top. You manually verify closely the unsafe parts, and if you got that short part right you can use the safe wrapper wherever you want safely.
That does not seem to be how it is implemented now. Try the example use-after-free on https://uptrace.dev/blog/posts/go-memory-arena.html under 'Address Sanitizer'. There is no error at all when you dereference, it silently keeps working as freed pointers often do (until they don't). Maybe they will add the virtual address space thing later.
Edit: the issue (https://github.com/golang/go/issues/51317) in which that comment was made also has some people referring to the arena being freed "when the runtime gets around to it" rather than immediately. I can't imagine why the virtual address space wouldn't be made poisonous inside the arena.Free method, but it's possible they do do it later. So your pointers maybe do stop working at an indeterminate time. Sounds pretty much like C. I even tried adding a few `runtime.GC()` calls before dereferencing but to no avail, couldn't get it to crash.
The Java 20 implementation has no such issues (https://gist.github.com/cormacrelf/8ddd3cc1b086e4ade93c029a9...), partly because there is no API for allocating a generic T in the arena, so all dereferences are through the managed MemorySegment API. It only seems to offer raw memory, which you can use to implement arrays of unboxed C-style structs, which is great for FFI and network buffers etc. This would be a good tradeoff for Go as well, surely.
Unboxed C types in Go can be done with cgo (it's not necessarily pleasant, but it works).
Such as? Java buffers use byte arrays by default as underlying storage.
Look for "Direct vs. non-direct buffers". Some buffers are backed by arrays, some are not.
You're basically deciding whether you are okay with a higher cost for I/O or a higher-cost for interacting with the data in the buffer.
Also, I fail to get your conclusion — direct buffers only have a higher cost for initialization/dealloc otherwise reading/writing should be as cheap as it gets. Byte buffers on the other hand do get write barriers but that is also unlikely to be hit that often (and I believe larger byte arrays get allocated in a separate region that doesn’t get moved).
Moreover, when the go team accomplishes the same goals by extending the language instead of using a library, this is well received.
I'm sorry, I'm not really sure what difference between Buffer and byte[] in Java I'm supposed to care about, except for the details of how these objects are allocated in memory. There is not really any semantic difference between them, as far as I can tell.
Maybe I don't understand what you're saying--I'd love for you to clarify the reasons why you think Go developers are, I guess, a little bit dense, or foolish, or whatever negative adjectives you like to apply to people who use Go (you know, rather than talking about the language itself--it's always fun to make fun of people).
Go GC does not move objects in memory (mark and sweep) while other GCs (mark and compact) do, to avoid heap fragmentation and better throughput (see generational GCs).
One issue with mark and compact GCs is that you can not move the array of bytes during a system call. A ByteBuffer/Memory object, is an object that allocates its byte arrays using malloc, outside of the heap, so the bytes are not moved by the GC.
What you're encountering is one of two exceptions in the implementation where you might not get an immediate failure:
1. If you use only a very small amount of an arena chunk and free it, it goes back on a reuse list as an optimization. Accessing that chunk's memory, despite the fact that the arena was freed, is entirely memory safe: nothing else will use that memory. That address space will be properly poisoned once the arena chunk is full, or close to it. 2. If the GC is actively marking, arena chunk poisoning is delayed to avoid races with the GC that might cause it to dereference a pointer into poisoned memory. The arena chunk is poisoned as soon as the GC is done.
The API explicitly does not guarantee a crash on use-after-free[1] because the Go team wanted a valid implementation of arenas to be simply "new" for New and noop for Free. The point is to just stay memory safe and have a high probability of catching an issue _in production_ where presumably arenas are filled up (otherwise why are you using arenas?).
Edit: arguably that optimization should just be turned off for MSAN/ASAN mode for greater user-friendliness, which seems reasonable. I think that was just an oversight.
[1]: https://cs.opensource.google/go/go/+/master:src/arena/arena....
I still have a pointer into the chunk. If you reuse the chunk in a different arena, I still have a pointer into the chunk. At no point is it invalidated, and the new arena will now start allocating new stuff from the start of the chunk. Right? And my old pointer still works, and the data inside it is at some point overwritten. How is that memory safe?
Are you eventually unmapping the virtual address space and remapping the underlying allocation somewhere else? So the poisoning happens when the chunk is picked up for reuse by a new arena?
It doesn't allocate at the start of the chunk, it just picks up wherever the last one left off. That allocated memory in the chunk from previous arena allocations is not reused until the chunk as a whole is unmapped and the GC can confirm that no more pointers point into that chunk's address space (if you leave a dangling pointer, you only waste address space). This points-into property is cheap to check, it's equivalent to whether the chunk has been marked by the GC.
Again, I think that MSAN/ASAN should probably just be more strict with these kinds of use-after-frees. You won't crash, but it's still technically incorrect. (Not much can be done about the "GC is in the mark phase" case, unfortunately. Otherwise MSAN/ASAN will complain when the GC inevitably tries to access a pointer into a delayed chunk.)
It would be nice for some debug/testing mechanism (akin to -race, or zig's detectLeaks) to proactively verify nothing has outlived the arena.
Wait until you learn how Rust now has a crate that implements seamless concurrent tracing GC, just like Go... https://redvice.org/2023/samsara-garbage-collector/
But if you have a compacting GC, you shouldn't need an arena. You should be able to get away with fixing the GC.
Practically speaking, "fixing the GC" is just a really, super hard problem. You're making a tradeoff between memory usage, CPU usage, pause times, and the performance of various features like pinned objects.
It makes sense that you may want different tradeoffs in different parts of your program. "Fixing the GC" is right up there with "just make a smart enough compiler" type sentiment that gets you in trouble--it's easy to say that you desire a better GC or better compiler, but when you actually try to make a better GC or better compiler, you find out that it doesn't solve the problems you were hoping it would solve.
Go's GC is tuned aggressively for short pause times. A lot of people actually want short pause times in their GCs, enough so that it's a selling point for Go. In Go, those short pause times were achieved partly by sacrificing CPU efficiency. You can't just go in and "fix" the CPU efficiency problems, but you can make new APIs that give you an escape hatch.
Adding more unsafe tools to Go or Rust do not diminish their safety guarantees. In fact, when Go and Rust add new unsafe features, they generally do so in a fashion that is much more constrained and has better defined edges than traditional methods, so that it's easier to use them correctly and easier to detect misuses. For example, the borrow checker still runs in Rust unsafe blocks; you still have to circumvent the borrow checker even once you're in unsafe. As for this arena package, it seems to have a few safety mechanisms, and hell, I don't even think it's off the table that it could be made entirely memory safe. This is especially true since memory safe does not mean "free of runtime errors", so all that is needed is to be able to detect the error conditions without allowing potentially undefined behavior to occur first.
Programming language experiments like Go's new arena package are good; they allow exploring what you can do to improve old concepts.
of course, this is completely wrong, for anyone reading this. that is simply not how systems work
The unsafe package in Go provides unsafe.Pointer, which allows you to convert between a pointer and integral type, essentially allowing you to make pointers of any type to any address, giving up type safety. Without it, Go pointers are safe, because the GC will never free memory that has live pointers, and you can't do math on Go pointers.
Yes, it's because Go's GC is primarily tuned for latency at the cost of (memory allocation) throughput. Furthermore, there are only a few knobs for Go's GC.
If Go's GC is implemented a bit differently and has a few more knobs so that it can be optimized for memory allocation throughput, maybe Go doesn't need an arena library. I don't even know other GC-ed languages that have/need an arena library because their GC is customizable enough.
Another option is https://pkg.go.dev/lukechampine.com/freeze -- not exactly what you want, but if the object is mutated after freezing, it will panic.
Is a shared err type package on the way out? I use it to bubble up HTTP status codes consistently, is there a better way? Should you always use sentinel errors?
Then, it was a breath of fresh air. Nowadays.. I find myself sighing. It works but the joy has faded.
Bit rot is life.
And then, slowly over a couple of decades, other people will be making it their life's work to add in the things that you thought were stupid because inevitably the need for them surfaces.
Hell, even C underwent pretty major changes from 72 onwards, because every compiler supported entirely different features. Things like void functions and returning structs or unions. Granted, this predated internet distribution of software, so changes overall were perhaps slower within a single implementation, but there were still radical developments happening on a frequent basis.
Only if a language with basically zero new concepts could have just learned from the litany of other managed languages’ mistakes..
So I keep writing my old ways code in Go and act like the new features don't exist.
Big sigh
Except that now if someone does `errors.Join` and they pass it to existing code that was using `errors.Unwrap` to inspect an error chain...
... they now get a not-unwrap-able error. Which they can still `As` to inspect...
...but since it's a tree, they can't recursively-`As` to find all instances of a type of error in a chain, like they could before (if you find something in one branch, you can only traverse that branch, because As doesn't maintain iteration-state).
It's not an unambiguous win, sadly. New code interacting with old code might misbehave.
For behavior: sometimes wrapping order matters, and As is convoluted to use to determine that, to say the least. Manually unwrapping is easy, and reasonably safe and easy - the standard library does it! But...
I think shared types from a package make a ton of sense and are really practical for writing flexible and maintainable code.