File System Interfaces for Go – Draft Design
go.googlesource.com
go.googlesource.com
Network file system could use the same interfaces + timeout settings.
If you can't tell, I don't like context. I've said before [0] I really hope that Go 2 comes up with an actual solution to the cancellation and task-local storage problems and deprecates context. Some comments in that thread pointed to alternatives that looked pretty decent, I wonder what state they're in these days.
It needs to 'infect' (as you put it) everything because it needs to pervade everything, so a library can pass it through to anything it calls. This is a good argument for making it implicit.
> And it includes an untyped key-value store implemented as a linked-list of pairs (!!!!), because why not?
That is just a cons chain, famous for example as the foundation of the Lisp programming language, and there is nothing wrong with it. Note that it is unfair to say that it is untyped: the values themselves are strongly typed, and the cells too are typed — interface{} (or T in Lisp terms) is still a type!
A cons chain has advantages for inheritance of values in a DAG of calls. It is not necessarily the most efficient, but simplicity is a virtue.
0: Continuation-passing style is both really powerful and well-nigh unreadable, for a reason.
I really hope they'll stay.
The term is a mistake, as it creates the impression that a backwards-incompatible product is in the works.
I wonder what poor abstraction the Go authors will come up with instead?
Every "monadic solution" prints the same code block without explaining how it would work, the types of the various variables, the semantics of the <- operator. I didn't leave the page with an understanding of how monads achieve these tasks.
This is apparently the monads curse where people that get them can't effing explain them to save their life.
The issue is pretty much a language barrier. All these articles talking about benefits of / interesting ways to use monads are written by people who speak the language, assuming the reader speaks the language as well. As with many functional programming topics, the fundamentals aren't incredibly easy to wrap your head around. But if you already understood monads, can you imagine how annoying it would be if every resource, discussion, article, etc. relating to them started with a pages-long introduction on What is a Monad?
In Haskell, this is a fundamental topic. Monads are used everywhere. If people always explained how they worked when discussing them, it would be like looking up sorting algorithms and having every algorithm description start with a long-winded explanation of how for loops work.
I agree that it would be unreasonable for every article that uses monads to describe what they are. But I don't think it would be unreasonable for every one of them to link to another article that does explain them.
That said, I can try to explain here:
Haskell has a bit of syntax available called "do notation". You can write Haskell without it, but it makes some things read better (as a matter of common but not universal opinion).
There's a simple, purely syntactic translation from do notation into regular application of functions. "Syntactic sugar causes cancer of the semicolon." There are four rules, none of which is complicated, and only two of which are relevant here:
First, a single expression is just that expression, nothing magic happens.
do
expr
simply becomes expr
Next, the arrows: do
x <- m
... more stuff, which might use x ...
becomes bind m (\ x -> do
... more stuff, which might use x ...
)
or in a few other syntaxes: bind(m, x => do ... more stuff, which might use x ...)
(bind m (lambda (x) do ... more stuff, which might use x ...)
m.bind(|x| { do ... more stuff, which might use x ...)
bind(m, [] (auto x) { do ... more stuff, which might use x ...)
That internal do is then expanded recursively.So to translate the whole block that's repeated throughout the article:
do
a <- getData
b <- getMoreData a
c <- getMoreData b
d <- getEvenMoreData a c
print d
becomes bind getData (\ a -> do
b <- getMoreData a
c <- getMoreData b
d <- getEvenMoreData a c
print d)
which becomes bind getData (\ a ->
bind (getMoreData a) (\ b -> do
c <- getMoreData b
d <- getEvenMoreData a c
print d))
which then becomes: bind getData (\ a ->
bind (getMoreData a) (\ b ->
bind (getMoreData b) (\ c -> do
d <- getEvenMoreData a c
print d)))
and then: bind getData (\ a ->
bind (getMoreData a) (\ b ->
bind (getMoreData b) (\ c ->
bind (getEvenMoreData a c) -> (\ d -> do
print d))))
and finally bind getData (\ a ->
bind (getMoreData a) (\ b ->
bind (getMoreData b) (\ c ->
bind (getEvenMoreData a c) -> (\ d ->
print d))))
Which is "just" a bunch of chained functions combining lambdas.So... why bother? and how does it do so many different things? and how does it know which to do? and what does it even do?
The key is that we're overloading "bind", picking the behavior that we want. You could do this in most languages by passing in a choice of function. We could have "do" take a parameter, like
do(bind)
or in an OO language you might hang bind on the objects involved. We often see this for particular instances - for instance, .then for promises/futures.Haskell does it with a mechanism called "type classes", where you can specify an implementation of an interface for a given type, and the compiler will figure out which implementation to provide. This is very similar to Traits in rust, implicits in Scala, etc. You can usually avoid specifying the types manually because of inference.
So in Haskell, Monad is an interface providing two functions:
class Monad m where
bind :: m a -> (a -> m b) -> m b
pure :: a -> m a
We've already come across `bind`, which lets us operate "inside" a context in a way that combines the contexts.The other function included, `pure`, takes a value and gives it the minimal possible context. What that means is implied by the "monad laws", which tell us how `bind` and `pure` must interact.
For each type that "is a monad", we tell the Haskell compiler how that type implements the interface:
// optionality:
instance Monad Maybe where
bind Nothing _ = Nothing
bind (Just x) f = f x
pure x = Just x
// state:
newtype State s a = State (s -> (a, s))
instance Monad State where
// pure gives the State action that produces x and leaves the state unchanged
pure x = State (\ s -> (x, s))
// bind threads the state through the chained actions
bind (State f0) f1 = State (\ s ->
let (x, s') = f0 s
State f2 = f1 x
in f2 s')
// lists:
instance Monad [] where
pure x = [x]
// bind for lists is concatMap
bind [] _ = []
bind (x:xs) f = f x ++ bind xs f
... that got long. Please ask any questions you're left with :) newtype Reader r a = Reader { runReader :: r -> a }
instance Monad (Reader r) where
pure x = Reader (\ _ -> x)
m >>= f = Reader (\ r -> runReader (f (runReader m r)) r)
which allows us to define: getContext :: Reader r r
getContext = Reader (\ r -> r)
and which spares us threading the context manually.I don't think that's such a big deal, so long as there's only one thing we're threading.
In some languages, we can be generic over "contexts which provide the thing I want, whether they do other things or not", which is sometimes a much bigger win.
The Maybe monad deals with the same issue. You have a series of function calls, each of which might or might not be legal, because you never know if you actually have the parameter for the next function call. Maybe you do, maybe you don't.
If you take the naive approach to solving the issue, you pass along Schrödinger's Cat each step along the way. You intertwine the concern of the algorithm you're trying to write with the uncertainty of whether you're carrying a live cat or a dead one. It can work, but it's ugly.
The monadic approach allows you to separate these concerns. You write your algorithm as if you know for a fact that the cat is alive. If the cat were ever to be revealed to be dead along the path, it doesn't matter, the monad separated the consequences of dealing with the live or dead cat from the rest of your algorithm. The rest of your algorithm simply doesn't get run.
Context is the same. You get to write code as if the context is always valid. You don't have to worry about context being cancelled, or timed out, or anything else. You take all those concerns and relate to them separately, in one place, where they can be neatly dealt with.
The whole point of the monadic pattern is to propagate state in such a way that it doesn't interfere with the pure algorithm which you're trying to write. You write the pure algorithm separately, and then use it within the monadic context of Context, so to speak.
func (f SomeConcreteFS) WithContext(ctx context.Context) fs.FS
could hold a reference to the underlying filesystem and call appropriate read deadlines, timeouts, etc. var fs fs.Fs
fs = somefs.New()
fs = contextfs.New(context.TODO(), fs)
fs is now an fs with an imbedded context.Looks like they mention that package in their proposal.
https://www.reddit.com/r/golang/comments/hv976o/qa_iofs_draf...
It was prompted by this proposal for os.Readdirentries():
https://github.com/golang/go/issues/40352
It seems to me that FileInfo isn't a suitable interface for either a general filesystem construct, or a performant implementation.
The problem is the exact details of what file attributes are supported vary widely from system to system.
The best approach is to support a subset which is reasonably common across platforms – file type[1], modification timestamps, file size in bytes, etc – and an extension mechanism to enable platform-specific attributes.
Java NIO handles this reasonably well with the java.nio.file.attribute package[2] in my opinion. (Not sure how easy it would be to port the concepts of that to Go though.)
IANA has a registry of OS-specific facts (i.e. file attributes) and OS-specific file types[3] – this is for use of FTP MLST and MLSD commands[4] but the registry is rather empty because that RFC doesn't appear to have got much adoption. It is a good idea though.
[1] There is a standard list of file types most platforms support – regular file, directory, link – but there are lots of special file types specific to various platforms (e.g. named pipes, UNIX domain sockets, BSD whiteouts, NTFS junctions), plus some filesystems have different subtypes of regular files or directories. (For example, on IBM z/OS, a "regular file" could be a UNIX file, a VSAM dataset, or a non-VSAM dataset, and the later two both have several subtypes; similarly, z/OS has UNIX directories, but PDS(E) could also be viewed as a non-UNIX type of directory.)
[2] https://docs.oracle.com/en/java/javase/14/docs/api/java.base...
[3] https://www.iana.org/assignments/os-specific-parameters/os-s...
https://www.reddit.com/r/golang/comments/hv976o/qa_iofs_draf...
Reading the whole dir into an array is bad enough as it is, but they even call stat() on every file too.
If an API is going to arbitrarily sort filenames, at least let us pass in a comparator function so we have some control. That would actually be a decent improvement to its API.
While I can't speak for Go's authors' mindset, here's my view:
One array of records is about as simple interface as possible, and it's good enough for many use cases. Contrasting with that, reading the directory with repeated readdir()'s or equivalent is fraught with hidden pitfalls. You end up being responsible for sorting, for resuming syscall in case of a signal or other interruption, for handling changes to the list of files during read... and probably other concerns above and beyond that.
os.Readdir() fetched 10.000 file names into a slice in 0.2s.
That's more than fine for most use cases and certainly doesn't look like "one of the stupidest, most broken interfaces I've ever seen".
Specially given that it's not the only way to iterate over files of a directory in Go. There's filepath.Walk() for one.
> The files are walked in lexical order, which makes the output deterministic but means that for very large directories Walk can be inefficient.
If it's fine for "most use cases," that's all well and good. Just implement the correct interface, which is fine for ALL use cases, and then implement this sugar on top of that.
Seems like they postponed incorporating fastwalk into the standard library because it would break API: https://godoc.org/golang.org/x/tools/internal/fastwalk
Perhaps until then one could use libraries that expose fastwalk:
I now built the exe and ran:
real 0m0.199s
user 0m0.000s
sys 0m0.014s
So 0.2s for 10k files on the worst possible hardware/software scenario I could find nearby. Edited my original comment. Thanks!It is still a ton of time, about an order of magnitude more than optimal if the real/sys time split is to be believed.
It's not a "stupid" interface if it does exactly what you need. And I think this is a pretty common use case.
Otherwise, you'd better be 110% sure of this:
> Even when you know that you're dealing with a small number of files, and performance isn't critical?
Because if you turn out to be wrong down the line, your only option is now to rewrite the whole thing in a less-stupid language. Might as well just save yourself the trouble and do that from the beginning.
just use it when you need it and be aware of the downsides...
Creating an API where the entire set of data is returned from a data source, without an option to limit the result, is bad. It means that, if I have to read a directory with 2000 files (and this is quite possible! - think /bin, etc.) the function will have to allocate a slice with 2000 entries, and then call stat on all of them.
Stat is not free, especially on certain filesystems (NFS!), nor is reading directory entries.
It's a mega convenient function if you want to slurp in a few files.
What does being aware of the downsides get you?
Is there some alternate interface that avoids the downsides?
You opendir, read 100 entries, do whatever with them, read 100 more, repeat, then closedir. So it's a buffered interface similar to normal I/O.
There's been proposals for doing readdir+stat in one syscall, but nothing merged yet. https://lwn.net/Articles/606995/
Reg downsides: Maybe don't use it on a directory with 1000 entries...? Its good to be aware of that or not?
It’s fine to include an inefficient version if you also have a ‘proper’ way to fall back to when performance does matter. And if performance does not matter, why are you even using Go? Funnily enough Python Ruby and Node do expose performant ways to do this in the stdlib, so...
I think they named the call after the posix one. If so, they must have known about opendir/closedir.
Also, thanks for the ad hominem, perhaps you shouldn’t attach your self worth to a programming language and see how that works out for you.
Abfs is conceptually identical to the go draft design, including the concept of "extension interfaces", although at the fs interface level not the file interface level. One of the problems I ran into early was the need for an easy to use file handle. The `net/http` FileSystem interface shows that you must always implement a custom Open function and a custom File. An object that implements `net/http` File is not adequate, it must actually be cast to http.File. Because of this I've found that it is burdensome to allow flexibility in the definition of the File interface because on the one hand specific functionality not obviously available to a generic File interface must be looked for using type assertion and the consequences of not finding it are poorly defined (should we return an error, should we proceed with some work around, etc), and on the other functions that return a file handle have many possible choices making un-anticipated interoperability unlikely. So instead I opted for always returning a `absfs/absfs` File interface, but return a ErrNotImplemented for functions that are not supported by an implementation. I don't like this, but I like it better then having to do a lot of interface unwrapping to get to the same place, and having it be an error makes the consequences more concrete.
Nevertheless, I'm a million percent behind having a filesystem abstraction in the go standard library. It is immediately useful in testing to redirect potentially costly io operations. It is composable, allowing you to do very little work to wrap a file system with mutexes, timouts, and other gating mechanisms. It allows you to support a transactional file system by spawning a FileSystem from another FileSystem. Caching, copy on write, layering, and arbitrary data transformations are all much easier to reason about and implement using a prototypical filesystem as the model. Cheers, thanks for this Russ and Rob! From where I'm standing it would be one of the most valuable improvements to the go ecosystem possible at this point!
Instead of a language with more features than one can count, we have something really concise, yet powerful. This is something I miss a lot about Go and I think can perhaps even have a greater impact than generics for most people (at least I feel like this is my case), yet this is only being considered now after ideas matured in the community and people developed ways around it (see https://github.com/golang/go/issues/35950).
The outcome will probably be something that will be robust and stable and once in the language will likely last a long time unchanged without feeling awkward.
This blog post is just one example of those kinds of issues, and it's just the most recent of such that I've read: https://fasterthanli.me/articles/i-want-off-mr-golangs-wild-...
Things like no package management, no vendoring, importing modules directly from GitHub without any concept of versioning, the confusing mess that is $GOPATH, the awkward handling of errors, and so on. Lots of things that were fixed, but shouldn't have been broken in the first place.
Go certainly has use and functionality and benefits (goroutines sound awesome), but Go itself seems like kind of an awkward mess when viewed outside of the context of "internal Google tool".
* no package management
I use Go Modules for Pion and really love them https://github.com/pion/webrtc/blob/master/go.mod I select the versions I want and everyone that downloads my code uses the proper version.
* No vendoring
Just do `go mod vendor` and you are done
* importing modules directly from GitHub without any concept of versioning
You set the versions in your `go.mod` file.
> Go itself seems like kind of an awkward mess when viewed outside of the context of "internal Google tool".
Go is actually really nice from a community standpoint. WebRTC is driven by Google and is magnitudes worse. I deal with a wonderful mix of arrogance and ignorance. I appreciate the great tech they bankroll and try not to let the other stuff bother me.
That the things needed fixing (and especially in cases where there were previous ways of doing things that had to be entirely discarded) seems to go against the idea that there was thinking ahead in those areas.
But maybe you are right and they didn't anticipate these issues! I have no idea either way :)
Probably doesn't help that at google when a package breaks they'll just force some faceless drone to make it work. And since they don't have customers in the traditional sense that works for them.
You just said "no dependency management" 4 times in a row.
go is mega awesome and I love going back to the projects where I use it. It's a simple language that makes it very easy to write good code. And the library ecosystem is pretty good as well, especially when it comes to networking stuff.
The Go authors have explicitly stated they don't do things until they figure out the right way to do them. No language gets it all right out of the gate, but Go got pretty close with e.g. its standard library.
Come on.
He should try golang and then come back instead of relying on blogposts to form his opinion
What they are not stating is that Go is a conservative language with a standard library hampered by that conservatism, nor are they making a judgment on the current state of the language.
If you are going to criticize someone's post, give them the respect of actually reading their argument, and not just writing a knee-jerk reaction.
You can read it https://slashdot.org/story/11794
He didn’t even need a time machine.
In fact, in comparison, Go is pretty unstable compared to those two once you take into account the total time they have existed.
It is definitely not unique to Go.
I'm sorry to say but I like PHP better than go and php is in the top 5 of languages people use that I hate (most to least is go, rust, js, java, php). You might say they're the most popular languages but so my top 3 languages. From least to most: D, Zig, C++, Python (3) and C#
There is no exception for this API either.
I kind of feel weird when there has been loud praise on Go team, while in general it was no mention of the decade of hard work of evolving C++ APIs...
Like you I read this proposal through that lens, but I actually don't see a line from Google's File abstraction to this. Is this Go abstraction really forward-compatible with cloud-native filesystems? And when I say forward compatible I mean would it work with Google's decade-old Colossus? I think it isn't because it puts Stat in the required interface and Stat returns FileInfo, many of the fields of which might be meaningless in a cloud filesystem (like mode, which is a unixism, and IsDir, which doesn't make sense for non-hierarchical filesystems).
If you can't tell I'm not much of a fan of trying to abstract over filesystems. Mostly they don't resemble each other at all, except in some extremely high level concepts.
Then, when I saw people praises Go APIs, I feel that some people worked on C++ APIs at Google, were missing some credits. That's a derived feeling from the above.
That is not to say that this File design was already being praised (not to say that it does not deserve praise); or that it should be criticized (not to say that it does not have problems).
As for your example to contradict my claim: "Stat returns FileInfo, many of the fields of which might be meaningless in a cloud filesystem": 1. This has to be there because a File API has to be compatible with OS files (Unixy). And Google's API also does the same thing. 2. Incompatibility derives from enforcing information that not present in another scenario. Clearly, the presence of such attributes does not prevent them to be applied on cloud files systems.
And further, Cloud file systems do have mode, and directory...
The most popular cloud file system, S3, doesn't have folders. It has prefixes, but if you treat them like folders, you're in for a lot of pain.
http://www.pdsw.org/pdsw-discs17/slides/PDSW-DISCS-Google-Ke...
Page 9
I am not familiar with S3.
In terms of the ideal of preferring types as a means to constrain functionality, I think FileInfo fields should have been expressed as granularly as possible using extension interfaces in this case, but practically I'd prefer putting some placeholder values in FileInfo for missing functionality over the clunky way these extensions have to be implemented in in Go.