Building a Streaming Platform in Go for Postgres
blog.peerdb.io
blog.peerdb.io
You also have to be very careful about managing the channel lifecycle. If you're not pulling (selecting from) the channel, the library will be permanently stuck. So you must now have a way to tell the library to stop sending, and it must cancel any in-flight send operations if you call producer.Stop() or whatever. In my experience libraries often have bugs in their channel code. It's far too easy to get deadlocks with channels that have interdependencies, and you have to be very careful about buffered versus unbuffered channels, as they behave differently.
A better API, in my opinion, is to offer a callback or single-method interface. Then the implementer of that callback or interface can choose to use channels internally if they desire, or they can use something else. You get the same backpressure support since you can treat it as synchronous.
After all, a channel's send interface is essentially just:
type Channel[T any] interface {
Send(T)
}
But a "chan T" doesn't offer this flexibility.My rule of thumb for channels is that they're goroutine glue, not an API primitive. Build APIs out of interfaces, not channels. The only thing that uses channels should be the one that's controlling the goroutines, because it's the thing that orchestrates them.
That said, it's not a hard rule. There are places where channels may have their place in a public API, though I'm not sure I can think of any examples off-hand.
you can always wrap channels to make them worse and less capable, but your API should expose the more capable option.
As you say being able to select is really nice too.
I found your excuse above is really nonsense. when your program is in Golang, you've already picked side, the concerned concurrency model has already been chosen by the user.
we are not talking about one of random concurrency models, we are talking about channel based sychnronization and communication in golang, if you don't want that and consider it as an issue, you shouldn't be using golang in the first place.
If I wanted to encapsulate iterating over a channel of Records, maybe it would be something like Go's io.Pipe function [2], which returns a PipeReader and PipeWriter? Except that it would work on Records rather than byte streams.
I don't have enough context to know if the extra encapsulation is a good idea in this case, though.
[1] https://github.com/search?q=repo%3APeerDB-io%2Fpeerdb%20GetR... [2] https://pkg.go.dev/io#Pipe
The whole purposes of having such built-in channel as a part of the core language is to encourage everyone to use it whenever possible, the reason is dead simple - it is much less likely to shoot yourself in the foot.
Done() <- chan struct{}
What are you trying to say? That using a chan as a function argument can shoot yourself in the foot? The whole idea behind channels is to not shoot yourself in the foot with mutex's, deadlocks, etc. What are you trying to say? Wrapping a chan in a type handle is safer than passing a chan of a type? Nope.I appreciate your concerns but one could say similar things about nil pointers or accessing slices out of range. Anything you use should be handled with care.
you can't abstract over channels easily unless you're missing the point of channels.
What I don't understand is why there should be such a blog article telling people that you decided to use channel in your go based program? This is like writing an article telling people that you decided to use the MMU when building an OS.
We added Queues as a supported target, where one of our users wanted single digit second latency. This is when we introduced channels. We agree that it is a small change. But the latency improvements were significant and wanted to share it with broader programming/Go community of Go Channels and the objective impact they can have.
All valuable details were omitted, e.g. 1) how much data is being moved? a few megabytes? what is the state of art latency? maybe 0.001 second for that? when latency is still seconds huge, what can be further improved, in the go runtime or in the apps? etc
A highly efficient pull/push model without using go channel can easily beat your 1-5 seconds latency. A crappy implementation using channel can get worse.
Delay of several seconds is when you pass on information mouth to mouth, in 2023, that kind of delay doesn't justify a blog article on such fancy implementation. I'd suggest you to delete the implementation to rebuild from ground zero.
We were observing 60<>40 ratio for Pull/Push, as specified in the blog and majority of the Push was target data-store (network) bound. So there was less room of optimization on the Go side.
The test was done with ~10-15K TPS on Postgres and the target was Azure Event Hubs, which has a limit of 1MB write batch size.
While it would certainly be possible to package them into individual binaries, I found it significantly easier to define the stack in a docker compose file with the requisite environment setup.