Go Slices Are Fat Pointers
nullprogram.com
nullprogram.com
https://dlang.org/spec/arrays.html#dynamic-arrays
My suggestion for C to adopt fat pointers:
So are Rust's.
Possibly more interestingly, so are its trait objects, the trait object pointer is a (vtable, instance) rather than (instance*, length).
Also of note: that applies to "owned pointers" (`Box`) not just "references", despite the former looking very struct-like: https://play.rust-lang.org/?version=stable&mode=debug&editio...
In a way, C++ and Rust (and possibly D?) have pretty much an infinity of fat pointers as anything can be made to deref' to something else, and be more than one pointer wide (e.g. C++'s std::string and std::vector and Rust's String and Vec)
> Possibly more interestingly, so are its trait objects, the trait object pointer is a (vtable,
> instance) rather than (instance*, length).
The Go interfaces are the same. See https://research.swtch.com/interfaces.https://security.stackexchange.com/questions/155572/how-are-...
Not that I'm against fat pointers in C, I always thought that in hindsight not having a primitive "slice" type was a mistake. In particular using NUL-terminated strings instead of slices is both error-prone and makes implementing even basic string-splitting functions like strtok annoying or inefficient (you need to be able to mutate the source string to insert the NULs or you have to allocate memory and copy each component. If you had slices you could easily return a portion of the string without mutation or copy).
In this case you get another problem: how to free memory when there might be many slices pointing at it.
extern void foo(char a[..]);
differ from extern void foo(size_t dim, char a[dim]);
if they are meant to be binary compatible?But anyway, you can use D in BetterC mode and have bounds checked slices!
(Jonathan Blow: Ideas about a new programming language for games.) https://www.youtube.com/watch?v=TH9VCN6UkyQ&list=PLmV5I2fxai...
When a language owns the platform its quality comes second.
#define COUNTOF(array) \
(sizeof(array) / sizeof(array[0]))
Generally, when you write a C macro, you should parenthesize each use of any argument inside the macro. The exceptions would be cases like token pasting where you want the literal argument text and only the text.I don't have a specific failure example in mind for the COUNTOF macro (maybe someone can post one?) - but here is a macro where it is easy to see the problem:
#define TIMESTWO( value ) ( value * 2 )
int ShouldBeEight = TIMESTWO( 2 + 2 ); // Oops, it's 6!
This happens because the macro call expands to: 2 + 2 * 2
Parentheses prevent this problem: #define TIMESTWO( value ) ( (value) * 2 )
The macro also illustrates a common misconception: sizeof is not a macro or function call, it's an operator. So wrapping the argument to sizeof in parens isn't necessary unless they are needed for other reasons. (They don't hurt anything, they just aren't needed.)So when I've written a COUNTOF macro (and it is a very handy macro in C/C++ code if your compiler doesn't already provide a similar one like __countof in MSVC), I write it like this:
#define COUNTOF( array ) \
( sizeof (array) / sizeof (array)[0] )
The space after each sizeof is my attempt to indicate that the following parens are not function/macro call parens, but are there specifically to wrap the array argument safely.There are various preprocessor tricks to avoid this stuff (eg “do{ x = (array); ... } while (0)”), but, honestly, if you want this sort of memory safety, C++ is your friend.
-Werror-sizeof-pointer-div
Or in C++: #include <type_traits>
template < typename T, std::size_t N >
constexpr std::size_t count_of(const T (&)[N]) { return N; }
#define COUNTOF(a) std::integral_constant<std::size_t, count_of(a)>::value
(C++03 friendly versions are entirely possible, but left as an exercise to the reader)The __countof macro in MSVC does go to some extra work to make it more safe, so it's definitely recommended to use that instead of rolling your own COUNTOF in that environment.
Even though array[0] is non-existent in this case, it's fine to pass it to sizeof to get the size of an array element as if there were one. After all, a pointer can point one element past the end of an array.
Hmm, I wouldn’t put it that way, since sizeof(array[1]) would work just as well. I like to think of sizeof running the type checker on the expression provided and then looking up the size of the type of the expression: this makes it clear that the thing inside can be completely bogus as long as it is syntactically correct.
String.substring in Java does exactly that. I'm sure numpy arrays in Python are implemented like that, perhaps slightly more fatter.
Yes, add one more level of indirection.
For instance traditional "oo" languages don't usually use fat pointers for dynamic dispatch, you have a single pointer to an instance which holds a pointer to its vtable.
In Rust however, the "object pointer" is a fat pointer of (vtable, instance). Go's interface pointers are the same (which combined with nils leads to the dreaded "typed nil" issue).
Secondly, I think that for a lot of people, when they have an interface that is itself nil, they mentally tag it with the label "nil", and when they have an interface that contains a typed value that is nil, they mentally tag this with "nil", and from the there the confusion is obvious. Mix in some concepts from other languages like C where nil instances are always invalid, so that you don't mentally have a model for a "legitimate instance of a type whose pointer is nil" [1], and it just gets worse.
Third, I think it's an action-at-a-distance problem. I'd say the point at which you have, say, an "io.Reader" value, and you put an invalid nil struct pointer in there that doesn't work, the problem is there. Values that are invalid in that way should never be created at all, so "interfaceVal == nil" should be all the check that is necessary. So you have a problem where the code that is blowing up trying to use the "invalid interface" is actually caused by an arbitrary-distant bit of code that created the value that shouldn't have been created in the first place, and that pattern generally causes problems with proper attribution of "fault" in code; the temptation to blame the place that crashed is very strong, and, I mean, that's generally a perfectly sensible default presumption so it's not like that's a crazy idea or anything.
[1]: For those who don't know, in Go, you can have a legitimate value of a pointer which is nil, and implements methods, because the "nil" still has a type the compiler/runtime tracks. I have a memory pool-type thing, for instance (different use case than sync.Pool, and predates it) which if it has an instance does its pooling thing, but if it is nil, falls back to just using make and letting the GC pick up the pieces.
I royally screwed up a deploy once because my functions were returning nil error implementations that were != nil and triggered failure paths.
It used to - it doesn’t these days, where it now copies.
Not since 2007 (java 7). The extra ints (8bytes) used for marking the beginning/end of the string and the fact the subscring leaked the ref. to the original string caused extra memory footprint. So now it's just a copy unless the substring is exactly the same as the original.
Previously (around 2004) I used to call new String(string.substring) in specific cases to prevent leaks.
>Is there any other way that can be done? <-- yes copies.
The absolute minimum meta info required for a multidim array is shape, stride. Numpy implements them as a [HUGENUMBER] array of integers for both. Gorgonia implements them as slices. Both of these incur a lot of additional movement of memory.
More here: https://youtu.be/fd4EPh2tYrk?t=654
https://en.wikipedia.org/wiki/Intel_MPX
as well as TinyCC's Bounds Check, https://bellard.org/tcc/tcc-doc.html#Bounds
which cites the aforementioned GCC patch.I think AddressSanitizer uses a shadow mapping structure more similar to Valgrind.
Although your point about performance still stands. Intel MPX is dead--support removed from GCC, IIRC, because of hints Intel was abandoning it. We'll never know how well it could have been optimized in practice. I think ultimately we simply need proper fat pointers, but those might not become ubiquitous (in the sense of becoming a de facto standard part of shared, C-based platform ABIs) until we see 128-bit native pointers.
Algol 68 slices are also fat pointers. Only they also carry a `stride` field, since in Algol 68 you can slice multidimensional arrays in any direction.
I don't think you could use OS APIs without boxing and unboxing data, but I think that's the case with every language except C.
One that I thought was especially nifty was [0].
[0] https://github.com/facebook/folly/blob/master/folly/PackedSy...
As a result of this you can use pointers as interface types without boxing them. For example you can take an interior pointer to an array, or a C pointer, and use it as an interface type without boxing since the type info doesn't need to be stored with the data.
C++ is a bit better with templates, but lacks type erasure (hence long compile times) and existential generics (so OOP isn't just syntactic sugar).
Is it that the Go documentation doesn't cover the underlying implementation at the machine word level?
uintptr addr = (uintptr_t)buf & 0xffffffffffff;
uintptr pack = (sizeof(buf) << 48) | p;
should be “| addr” #define ARRAYPTR(array) \
ADDROF(array, COUNTOF(array))
should be FATPTR(array, COUNTOF(array))> There is one important caveat: Go is not purely memory safe in the presence of concurrency. Sharing is legal and passing a pointer over a channel is idiomatic (and efficient).
> Some concurrency and functional programming experts are disappointed that Go does not take a write-once approach to value semantics in the context of concurrent computation, that Go is not more like Erlang for example. Again, the reason is largely about familiarity and suitability for the problem domain. Go's concurrent features work well in a context familiar to most programmers. Go enables simple, safe concurrent programming but does not forbid bad programming. We compensate by convention, training programmers to think about message passing as a version of ownership control. The motto is, "Don't communicate by sharing memory, share memory by communicating."
Every other language just looks awful.
But programming is fun and it’s especially fun to move easily among all these flawed tools and accomplish a goal.
Golang is ok, the runtime is pretty weak compared to modern JVMs and tight C++ is faster, Clojure and Haskell are denser. It’s ok but I wouldn’t stop checking out other stuff.
It's all about cost/benefit tradeoffs. Programming involves making cost/benefit decisions many times a day. Running a programming shop is that multiplied by the number of people, squared. Technical debt is about unresolved cost/benefit tradeoffs from the past.
Basically feature complete, with unitests, and interop with c arrays
A few changes for a later version:
* Pure public domain license (it is that already, but the wording could be better)
* Remove c++ support... Not sure about this
* Let higher order map/filter functions take a void* payload when using plain c functions over clang blocks (or lambdas)
Well, you can't dereference these fat pointers directly on amd64, you have to remove the tag.
aarch64 actually has a Top Byte Ignore flag though :)
I don't disagree with any of this article at a factual level, I'm just trying to understand the author's message.
I've now learned 2 things: what fat pointers are, and that go slices are fat pointer.
It was an interesting idea but you needed a whole new set of string functions and bridge functions. It takes a runtime (or codebase) wide commitment to make it work.
https://www.amazon.ca/Annotated-C-Reference-Manual/dp/020151...
You probably also wouldn’t want to use any string manipulation functions on it, but I’ve never had a reason to do that anyway.
I've personally resorted to using `unsafe` trickery in tight loops to overcome this limitation.
Also, no proper package manager.