Porting dl.google.com from C++ to Go
talks.golang.org
talks.golang.org
So he put the whole mess in a bin and re-done it cleanly with Go. Now it's much nicer. Some of Go's attributes helped along the way.
Did I miss something?
Though as the Node people occasionally point out, it is advantageous to have this sort of thing baked into the language, so that everything done in the language supports the concepts, rather than having a relatively small corner support it. Plus you get Go, instead of Java, which I for one would find an improvement.
The Java language specification does not require the use of OS threads for Java threads.
It all depends which JVM you refer to.
The JVMs that make use of green threads follow a similar threading model to Go.
Don't mix languages with implementations.
http://msdn.microsoft.com/en-us/library/vstudio/hh191443.asp...
TPL(Task parallel library) does use additional threads, and that's a different programming model.
The async/await model in C# is essentially from F#'s async workflows:
http://googleresearch.blogspot.com/2006/03/hiring-lake-wobeg...
What do the presentation (or our discussion) gain from the fact that Brad is a super smart guy "much more capable than the 'typical' programmer"? Why are we even bringing up that point? :)
As with all anecdotal language advocacy, you have to take it with a grain of salt, and in this case maybe more than usual.
Within Google today, if you're directly using local disk instead of the Google storage stack, you're going to have a bad time. Even more so if you're calling read() from a single-threaded event loop.
So, for me, that is somebody who didn't follow the state of Go libraries, it was much bigger new information in the presentation the author's certitude that some library covers everything of the standard (good to know how much care was invested there!) than the fact that when somebody has already made some very good tools he can more easily use them then probably anybody else who'd start without such a deep knowledge.
My own takeaway was "now that I know that at least some libraries are very well thought-out that's a good argument to consider when deciding about the use of Go in some future project." Still it's to weight against the learning curve and the quality and ease of use of the available libraries in other more common languages. In that light it's still good to know that the author of the presentation is really very competent.
Not directly related to your question, I also applaud the author's honesty in describing the problem he solved: first, the underlying assumptions changed: I can imagine that the main reason for fetching the data to the local disk in 2007 was that then fetching the data directly from the remote storage was much slower than it is now, making his current approach impossible then. Second, it's not that the original design was as inefficient as the later suboptimal modifications made it. Third, it was a project in which for a while nobody wanted to invest the time to. That's the convenient point to jump in and demonstrate what you can do with the new tools and your knowledge and time. Once you change the equation, a product once almost without the future can become a basis for future much more successful projects.
> The Go libraries are something very new. People who
> developed the libraries know them inside out. Outside of
> that group of people, there's certainly much less
> knowledge about them.
I guess it's a tautology that the people who developed something know it best, but you're being really disingenuous here by implying that only those people would be capable of doing a project like this. Go's stdlib HTTP server takes great pains to be accessible and powerful.So it's a word of confidence that, yes not only is the interface nice, but the implementation is also good (enough) -- and you probably won't regret leaning on it in a future project.
They know how to keep things simple
Fair enough, but it was a bit of a foregone conclusion that he wouldn't want to develop/reuse a C++ HTTP stack as he went into the project with the intention of using Go...
- Gofmt; code formatting
- Smaller number of language keywords compared to C++
- Garbage collection
- Built in concurrency
- Simpler error handling (no C++ style exceptions)
Could have been very similar if they'd rewritten it using C++11 features (the main difference being that you'd have to use a HTTP library instead of HTTP being in the standard library, but that doesn't make much difference)...
That said, I'll happily read this story over and over again, because it tends to provide insight in what works and what doesn't.
I guess that some details on the slides would answer those questions, but if anybody here know the answer (some slides were pretty obscure if you're not already familiar with file serving and/or google architecture).
0) copy all the files from the DFS to the local storage.
1) Attempt to make a distributed filesystem available as a mount point (for example, Hadoop FUSE) on the nginx server so it can serve the data.
2) Make the nginx server know the distributed filesystem API and talk to it via user space directly.
The former leads to insanity (writing a really good, fast FUSE implementation is devlishly hard). The latter can work; I've seen a number of open source codes adapted to distributed filesystems.
This is made easiest if the underlying includes an IO abstraction layer. From a quick skim, it seems nginx has a OS abstraction layer (win32 and unix implementations):, which is more than sufficient.: http://trac.nginx.org/nginx/browser/nginx/src/os/{win32,unix.... From that, it wouldn't be hard to write a "distributed FS variant" (although IO and OS are not orthogonal concepts).
The next problem you're going to have is performance. nginx and other systems really are written with the expectation that IO is local, high throughput, and low-latency. Opening a remote file often takes some time (hundreds of milliseconds), and serial roundtrips cause additional latency, especially if DFS and server are not within 5ms RTT.
The next logical step is to write a readahead layer (if you're serving files larger than a single read block to the DFS) to get better throughput.
A combination of the above techniques, with a bit of tuning and elbow grease, is a good foundation for scalable (in terms of file size, total # of files, etC) serving.
I think the biggest reason would be that almost any off-the-shelf server software would struggle in a Google server environment. They have solved scale at the machine level, and he alludes to this on one slide:
... why aren't you using the cluster file systems like everybody else?
... cluster file systems own disk time on your machine, not you.
> Does an off-the-shelf httpd support ACLs?
Not really. You inevitably end up writing custom code to conform to your particular requirements and/or existing systems. If you want high-performance, you end up writing it in C as a module for Apache/nginx/whatever.
> How easy would it be to make them support google storage instead of a file system?
Unless said storage system is presented to userspace through ordinary file interfaces, same as above. There's no general turn-key solution built into webservers. The problem space is too wide.
Using Go in this way gets you good performance, simple architecture, maintainability, and easy deployment with total flexibility to do whatever you need in order to solve your version of the problem. There are no straightjackets, you don't have to conform to (or find ways around) anyone else's conception of the problem space.
Surely, one would probably need a separate external tool that would pre-fill (nginx wouldn't know that some file's pending before it's requested) and clean up caches (provided that rules are complex than trivial LRU removal on some threshold), as one probably wouldn't like webserver doing this unsuitable job, but that should be another story.
Then again, I'm not sure why the caching wasn't better handled in the CDN that lives in front of it (cache hierarchies work well) leaving this server to simply serve only the very first request.
readable.pipe(writable);
Additionally the link to `http-proxy` on slide 30 is misleading; 60% of that file is comments, and about 50% of what's left is websocket support, with the rest being header parsing & redirect parsing. The actual proxying bit is very simple and straightforward, and if you don't need every feature `http-proxy` offers you can do it yourself with streams in < 10 lines.
As I mentioned in my talk, that code looks like fine Node.js code.
But it's still event-based, and the flow isn't readable. In the actual presentation I went through the code to show how control flow jumps around. I picked a Javascript project (and the top hit I got from a search) because people know Javascript.
Websocket support doesn't matter. In Go, you can also just io.Copy(websocket, src).
I agree Stream makes Node.js code better.
http://stackoverflow.com/questions/7479276/what-is-the-main-...
It's good for C#, but still a language wart that could be built-in. I like that Go only has one set of APIs for everything, not the sync way and the async way.
It's sad that C#, which started out as a fixed-up Java, is now growing its own warts.
Of course, Go's not perfect either.
I haven't yet tried Go, but I don't see how it could match the performance of C# API with a single function. C#'s async methods offload any IO to the process IO completion port threads, thus freeing the current thread to do more work.
Go generally uses synchronous functions, but a function in Go can be the subject of a "go" statement (sharing the name of the language should give an importance of how central this feature is), which causes the function to be executed as a goroutine (that is, asynchronously using an M:N threading model.)
If you enjoy CoffeeScript, Iced CoffeeScript does a great job of this too.
[1] http://alexeypetrushin.github.io/synchronize/docs/index.html
1 Executable programs or shell commands
2 System calls (functions provided by the kernel)
3 Library calls (functions within program libraries)
4 Special files (usually found in /dev)
5 File formats and conventions eg /etc/passwd
6 Games
7 Miscellaneous (including macro packages and
conventions), e.g. man(7), groff(7)
8 System administration commands (usually only for root)
9 Kernel routines [Non standard]"
So section 3 is a little wider than that, eg: $ apropos apache::xmlrpc
Apache::XMLRPC::Lite (3pm) - mod_perl-based
XML-RPC server with minimum configurationThis won't work so beautifully when whoever implemented all the magic bits under your business logic didn't do so to deliver reasonable performance for your use case.
Or if the underlying code is in fact wrong. This turns into the kind of code you can see from Line 242:
https://code.google.com/p/google-api-go-client/source/browse...
We optimise the standard library for the majority cases, and putting the download server into production helped us improve Go's HTTP stack for everyone. Similarly, when other users report issues with our standard library, we fix them.
The file you linked to contains some code to accommodate a bug that existed in Go 1.0 and was fixed in 1.1. Not sure what the relevance is here though.
So it leaves a bad taste when the presentation mocks the original C++ implementation while the Go solution only works because there is a fully functional high performance HTTP server already in the standard library.
There is not a single slide that tells us it may not be a good idea to rely on the standard library for such high-level functionality. Thats why I linked to the bug. A normal user has very little options when it turns out theres a bug in the standard library, other than to report it and wait for you to fix it. The problem is exacerbated since standard library and language versions are tightly coupled.
Go team is very welcome patches from the external contributors: http://golang.org/doc/contribute.html
No need to wait, you can fix it and have your patch accepted (after a code review).
After that, there's not such a difference between an HTTP server in the standard vs a third-party library
As for the license, it does matter since with the GPL you'd be stuck and it turns out that there's some debate with the LGPL as well (IANAL).
Anyway how would the GPL make a difference to your forking of the std lib? Either you link to the standard lib and all your code (that is linked into that binary and distributed as such) falls under the GPL, or you don't -- whether you distribute a patch (set) doesn't seem to change anything?
edit: And, if you don't give someone your patch, you don't have to give someone your patch?
Are you saying everyone should write their own HTTP server from scratch?
> cp -R $GOROOT/src/pkg/net/http ./net/http
Then simply change all "net/http" imports to "./net/http". Now you're free to carry out any changes you'd like.
P.S: Just because you can do this doesn't mean you should.
Also, it's a sign of a defective language that the library guards against old environments? Really?
Anybody has an alternative to read these slides? The content itself seems quite interesting
Edit: Scanned the source, looks a like a best-effort distributed lock, rather than any sort of consensus protocol. This works for a cache setting, where e.g. having a split-brain scenario and duplicating the work is no big deal.
My main gripe and reason for not jumping on the Go bandwagon is the error handling strategy, it was exactly what I wanted to see in the Go version.
But thanks for the info anyhow.
The Go code may look good in 2012 and 2013, while it's still fresh. But I'd be very curious to see how it looks in 2017 or 2018, assuming it's still even being used then.
[1]: https://groups.google.com/forum/#!searchin/golang-nuts/dl.go...
Some of the logic is ported from C++ to Go almost line-for-line.
Some of the architectural parts are completely redone.
But it has the same binary name and flags and RPC interface
func Copy(dst Writer, src Reader) (written int64, err error) {
// If the reader has a WriteTo method, use it to do the copy.
// Avoids an allocation and a copy.
if wt, ok := src.(WriterTo); ok {
return wt.WriteTo(dst)
}
// Similarly, if the writer has a ReadFrom method, use it to do the copy.
if rt, ok := dst.(ReaderFrom); ok {
return rt.ReadFrom(src)
}
buf := make([]byte, 32*1024)
for {
nr, er := src.Read(buf)
if nr > 0 {
nw, ew := dst.Write(buf[0:nr])
if nw > 0 {
written += int64(nw)
}
if ew != nil {
err = ew
break
}
if nr != nw {
err = ErrShortWrite
break
}
}
if er == EOF {
break
}
if er != nil {
err = er
break
}
}
return written, err
}
Uh, big deal?The chunk of what's important isn't explained at all:
- runtime/ takes care of memory management quite efficiently with a decent tracing gc in runtime/mgc0.c. I haven't benchmarked it against other stop-the-world collectors, but it should be no match for truly concurrent gc.
- runtime/proc.c schedules various blocking and non-blocking (called netpoll, which resolves to epoll on systems where it is available) calls. It seems to account for number of cores and use native threads, but I'm not sure how it interacts with the Linux scheduler.
- runtime/malloc.goc is the core memory allocator/deallocator. Seems to be a relatively straighforward arena allocator using a bitmap.
I didn't have time to go through groupcache, but the presentation certainly didn't tell me much about it.
Go is pleasant, but there are Go problems an Python problems (and C problems, Lisp problems, and so on)
We actually need more of this on HN, not less.
package main
import "fmt"
func main() {
cake := make([]byte, 2600)
copy(cake, []byte("Happy Birthday, Lundberg !"))
lundberg := cake[:8]
milton := cake[:0]
fmt.Println("Lundberg got :", string(lundberg))
fmt.Println("Milton got :", string(milton))
}
// :-) (Ran out of UTF8 smileys and movie fact checkers).
// And yes could've used string but bytes are awesomer.http://code.google.com/p/google-api-go-client/source/checkou...
But these slides were genuinely interesting to pretty much the whole HN audience: it covers architecture decisions, the effect of code rot, the breadth of the Go standard library, a new object caching system, and a real-world example of Go usage.
tl;dr even to someone who is a bit critical towards Go (ie. me), it was a very interesting read.