How fast are Linux pipes anyway?
mazzo.li
mazzo.li
However, none of the variants using vmsplice (i.e., all but the slowest) are safe. When you gift [1] pages to the kernel there is no reliable general purpose way to know when the pages are safe to reuse again.
This post (and the earlier FizzBuzz variant) try to get around this by assuming the pages are available again after "pipe size" bytes have been written after the gift, _but this is not true in general_. For example, the read side may also use splice-like calls to move the pages to another pipe or IO queue in zero-copy way so the lifetime of the page can extend beyond the original pipe.
This will show up as race conditions and spontaneously changing data where a downstream consumer sees the page suddenly change as it it overwritten by the original process.
The author of these splice methods, Jens Axboe, had proposed a mechanism which enabled you to determine when it was safe to reuse the page, but as far as I know nothing was ever merged. So the scenarios where you can use this are limited to those where you control both ends of the pipe and can be sure of the exact page lifetime.
---
[1] Specifically, using SPLICE_F_GIFT.
I haven't digested this comment fully yet, but just to be clear, I am _not_ using SPLICE_F_GIFT (and I don't think the fizzbuzz program is either). However I think what you're saying makes sense in general, SPLICE_F_GIFT or not.
Are you sure this unsafety depends on SPLICE_F_GIFT?
Also, do you have a reference to the discussions regarding this (presumably on LKML)?
If you don't use gift, you never know when the pages are free to use again, so in principle you need to keep writing to new buffers indefinitely. One "solution" to this problem is to gift the pages, in which case the kernel does the GC for you, but you need to churn through new pages constantly because you've gifted the old ones. Gift is especially useful when the page gifted can be used directly in the page cache (i.e., writing a file, not a pipe).
Without gift some consumption patterns may be safe but I think they are exactly those which involve a copy (not using gift means that a copy will occur for additional read-side scenarios). Ultimately the problem is that if some downstream process is able to get a zero-copy view of a page from an upstream writer, how can this be safe to concurrently modification? The pipe size trick is one way it could work, but it doesn't pan out because the pages may live beyond the immediately pipe (this is actually alluded in the FizzBuzz article where they mentioned things blew up if more than one pipe was involved).
I still think the man page of vmsplice is quite misleading! Specifically:
SPLICE_F_GIFT
The user pages are a gift to the kernel. The application may not modify
this memory ever, otherwise the page cache and on-disk data may differ.
Gifting pages to the kernel means that a subsequent splice(2)
SPLICE_F_MOVE can successfully move the pages; if this flag is not speci‐
fied, then a subsequent splice(2) SPLICE_F_MOVE must copy the pages.
Data must also be properly page aligned, both in memory and length.
To me, this indicates that if we're _not_ using SPLICE_F_GIFT downstream splices will be automatically taken care of, safety-wise.> This post (and the earlier FizzBuzz variant) try to get around this by assuming the pages are available again after "pipe size" bytes have been written after the gift, _but this is not true in general_. For example, the read side may also use splice-like calls to move the pages to another pipe or IO queue in zero-copy way so the lifetime of the page can extend beyond the original pipe.
The paragraph you quoted says that the "splice-like calls to move the pages" actually copy when SPLICE_F_GIFT is not specified. So perhaps the combination of not using SPLICE_F_GIFT and waiting until "pipe size" bytes have been written is safe.
Similarly for splicing from a pipe to a file or something like that. It's really the end(s) of the chain that want to (a) generate the data in memory or (b) read the data in memory that seem to create the problem.
I wonder if io_uring handles this (yet). io_uring is a newer async IO mechanism by the same author which tells you when your IOs have completed. So you might think it would:
* But from a quick look, I think its vmsplice equivalent operation just tells you when the syscall would have returned, so maybe not. [edit: actually, looks like there's not even an IORING_OP_VMSPLICE operation in the latest mainline tree yet, just drafts on lkml. Maybe if/when the vmsplice op is added, it will wait to return for the right time.]
* And in this case (no other syscalls or work to perform while waiting) I don't see any advantage in io_uring's read/write operations over just plain synchronous read/write.
That doesn't match what I've read. E.g. https://lwn.net/Articles/810414/ opens with "At its core, io_uring is a mechanism for performing asynchronous I/O, but it has been steadily growing beyond that use case and adding new capabilities."
More precisely:
* While most/all ops are async IO now, is there any reason to believe folks won't want to extend it to batch basically any hot-path non-vDSO syscall? As I said, batching doesn't help here, but it does in a lot of other scenarios.
* Several IORING_OP_s seem to be growing capabilities that aren't matched by like-named syscalls. E.g. IO without file descriptors, registered buffers, automatic buffer selection, multishot, and (as of a month ago) "ring mapped supplied buffers". Beyond the individual operation level, support for chains. Why not a mechanism that signals completion when the buffer passed to vmsplice is available for reuse? (Maybe by essentially delaying the vmsplice syscall'ss return [1], maybe by a second command, maybe by some extra completion event from the same command, details TBD.)
[1] edit: although I guess that's not ideal. The reader side could move the page and want to examine following bytes, but those won't get written until the writer sees the vmsplice return and issues further writes.
The vanilla io_uring fits "naturally" in an async model, but batching and some of the other capabilities it provide are definitely useful for stuff written to a synchronous model too.
Additionally, io_uring can avoid syscalls sometimes even without any explicit batching by the application, because it can poll the submission queue (root only, last time I checked unfortunately): so with the right setup a series of "synchronous" ops via io_uring (i.e., submit & immediately wait for the response) could happen with < 1 user-kernel transition per op, because the kernel is busy servicing ops directly from the incoming queue and the application gets the response during its polling phase before it waits.
---
[1] https://twitter.com/trav_downs/status/1532491167077572608
But from what I know about how vmsplice is implemented, gifting or not, it sounds like it should be unsafe anyhow.
https://mazzo.li/posts/fast-pipes.html#what-are-pipes-made-o...
I think the diagram near the start of this section has "head" and "tail" swapped.
Edit: Nevermind, I didn't read far enough.
None of those would impact the same physical page mapped into another process with its own V->P mapping.
That sounds like a security issue - the ability of an upstream generator process to write into the memory of a downstream reader process, or more perverser vice versa is even worser. I presume that the Linux kernel only lets this happen (zero copy) when the two processes are running as the same user?
But also, if you are sending data, why would you later read/process that send buffer?
The only attack vector I could imagine would be if one sender was splicing the same memory to two or more receivers. A malicious receiver with write access to the spliced memory could compromise other readers.
But that was way off and `seq` turned out to be ridiculously slow. I dug down a little and made a faster version of `seq`, that kind of got me what I wanted. But then noticed at the end that the point was moot anyway, because just piping it to the next program over the command line was going to be the slow point, so it didn't matter anyway.
@elysium pipetest % pipetest | pv > /dev/null
102GiB 0:00:13 [8.00GiB/s]
@elysium ~ % pv < /dev/zero > /dev/null
143GiB 0:00:04 [36.4GiB/s]
Not a valid comparison between the two machines because I don't know what the original machine is, but MacOS rarely comes out shining in this sort of comparison, and the simplistic approach here giving 8 GB/s rather than the author's 3.5 GB/s was better than I'd expected, even given the machine I'm using.(And as an aside, the combination of that font with the hand-drawn diagrams is really cool)
Here is the source: https://www.ibm.com/plex/
Hope that helps!
I was worried when I saw zfs send/receive used pipes for instance because of performance worries - but using it in reality I had no problems pushing 800MB/s+. It seemed limited by iop/s on my local disk arrays, not any limits in pipe performance.
More generally though, sidenote 5 says that the code in the article itself is incomplete and the real test code is available in the github repo: https://github.com/bitonic/pipes-speed-test
PHP comes in at about 900KiB/s:
php -r 'while (1) echo 1;' | pv > /dev/null
Python is about 50% faster at about 1.5MiB/s: python3 -c 'while (1): print (1, end="")' | pv > /dev/null
Javascript is slowest at around 200KiB/s: node -e 'while (1) process.stdout.write("1");' | pv > /dev/null
What's also interesting is that node crashes after about a minute: FATAL ERROR: Ineffective mark-compacts
near heap limit Allocation failed -
JavaScript heap out of memory
All results from within a Debian 10 docker container with the default repo versions of PHP, Python and Node.Update:
Checking with strace shows that Python caches the output:
strace python3 -c 'while (1): print (1, end="")' | pv > /dev/null
Outputs a series of: write(1, "11111111111111111111111111111111"..., 8193) = 8193
PHP and JS do not.So the Python equivalent would be:
python3 -c 'while (1): print (1, end="", flush=True)' | pv > /dev/null
Which makes it compareable to the speed of JS.Interesting, that PHP is over 4x faster than the Python and JS.
Python cheats, and it's still slow as heck even while cheating (buffers the output at 8192 chunks instead of issuing 1 byte writes).
write(1, "1", 1) loop in C pushes 6.38MiB/s on my PC. :)
It's not language "cheating" of course. It's just OP "measuring the wrong thing".
I get around 1.56MiB/s with that code. PHP gets 4.04MiB/s. Python gets 4.35MiB/s.
> What's also interesting is that node crashes after about a minute
I believe this is because `while(1)` runs so fast that there is no "idle" time for V8 to actually run GC. V8 is a strange beast, and this is just a guess of mine.
The following code shouldn't crash, give it a try:
node -e 'function write() {process.stdout.write("1"); process.nextTick(write)} write()' | pv > /dev/null
It's slower for me though, giving me 1.18MiB/s.More examples with Babashka and Clojure:
bb -e "(while true (print \"1\"))" | pv > /dev/null
513KiB/s clj -e "(while true (print \"1\"))" | pv > /dev/null
3.02MiB/s clj -e "(require '[clojure.java.io :refer [copy]]) (while true (copy \"1\" *out*))" | pv > /dev/null
3.53MiB/s clj -e "(while true (.println System/out \"1\"))" | pv > /dev/null
5.06MiB/sVersions: PHP 8.1.6, Python 3.10.4, NodeJS v18.3.0, Babashka v0.8.1, Clojure 1.11.1.1105
But using a longer string crashed after 23s here:
node -e 'function write() {process.stdout.write("1111111111222222222233333333334444444444555555555566666666667777777777888888888899999999990000000000"); process.nextTick(write)} write()' | pv > /dev/nullAlso, what NodeJS version are you on?
node is v10.24.0. (Default from the Debian 10 repo)
After some quick testing in a couple of versions, it seems like it got fixed in v11 at least (didn't test any minor/patch versions).
By the way, all versions up to NodeJS 12 (LTS) are "end of life", and should probably not be used if you're downloading 3rd party dependencies, as there are bunch of security fixes since then, that are not being backported.
Java has (had) weird idiosyncrasies like this as well, well it doesn't crash, but depending on the construct you can get performance degradations depending on how the language inserts safepoints (where the VM is at a knowable state and a thread can be safely paused for GC or whatever).
I don't know if this holds today, but I know there was a time where you basically wanted to avoid looping over long-type variables, as they had different semantics. The details are a bit fuzzy to me right now.
> I believe this is because `while(1)` runs so fast that there is no "idle" time for V8 to actually run GC. V8 is a strange beast, and this is just a guess of mine.
Not exactly: the GC is still running; it’s live memory that’s growing unbounded.
What’s going on here is that WritableStream is non-blocking; it has advisory backpressure, but if you ignore that it will do its best to accept writes anyway and keep them in a buffer until it can actually write them out. Since you’re not giving it any breathing room, that buffer just keeps growing until there’s no more memory left. `process.nextTick()` is presumably slowing things down enough on your system to give it a chance to drain the buffer. (I see there’s some discussion below about this changing by version; I’d guess that’s an artifact of other optimizations and such.)
To do this properly, you need to listen to the return value from `.write()` and, if it returns false, back off until the stream drains and there’s room in the buffer again.
Here’s the (not particularly optimized) function I use to do that:
async function writestream(chunks, stream) {
for await (const chunk of chunks) {
if (!stream.write(chunk)) {
// When write returns null, stream is starting to buffer and we need to wait for it to drain
// (otherwise we'll run out of memory!)
await new Promise(resolve => stream.once('drain', () => resolve()))
}
}
}
I do wish Node made it more obvious what was going on in this situation; this is a very common mistake with streams and it’s easy to not notice until things suddenly go very wrong.ETA: I should probably note that transform streams, `readable.pipe()`, `stream.pipeline()`, and the like all handle this stuff automatically. Here’s a one-liner, though it’s not especially fast:
node -e 'const {Readable} = require("stream"); Readable.from(function*(){while(1) yield "1"}()).pipe(process.stdout)' | pv > /dev/nullI'm reminded of this StackOverflow question, Why is reading lines from stdin much slower in C++ than Python?
Rust: 21.9MiB/s
Bash: 282KiB/s
PHP: 2.35MiB/s
Python: 2.30MiB/s
Node: 943KiB/s
In my case, node did not crash after about two minutes. I find it interesting that PHP and Python are comparable for me but not you, but I'm sure there's a plethora of reasons to explain that. I'm not surprised rust is vastly faster and bash vastly slower, I just thought it interesting to compare since I use those languages a lot.
Rust:
fn main() {
loop {
print!("1");
}
}
Bash (no discernible difference between echo and printf): while :; do printf "1"; done | pv > /dev/nullYou could probably answer this by replacing printf with /bin/echo and comparing the results. I'm not in front of a Linux box, or I'd try.
[0] https://www.gnu.org/software/bash/manual/html_node/Bash-Buil...
Ah, yeah, good point, I am wrong.
Interestingly, Go is very fast with a 8 KiB buffer (same as Python's), I get 218 MiB/s.
$ ./a.out 1000000 2000 | cat >/dev/null
buffer size: 1000000, num syscalls: 2000, perf:1578.779593 MiB/s
$ ./a.out 1 2000000 | cat >/dev/null
buffer size: 1, num syscalls: 2000000, perf:0.832587 MiB/s
Code is: #include <cstddef>
#include <random>
#include <chrono>
#include <cassert>
#include <array>
#include <cstdio>
#include <unistd.h>
#include <cstring>
#include <cstdlib>
int main(int argc, char **argv) {
int rv;
assert(argc == 3);
const unsigned int n = std::atoi(argv[1]);
char *buf = new char[n];
std::memset(buf, '1', n);
const unsigned int k = std::atoi(argv[2]);
auto start = std::chrono::high_resolution_clock::now();
for (size_t i = 0; i < k; i++) {
rv = write(1, buf, n);
assert(rv == int(n));
}
auto stop = std::chrono::high_resolution_clock::now();
auto duration = stop - start;
std::chrono::duration<double> secs = duration;
std::fprintf(stderr, "buffer size: %d, num syscalls: %d, perf:%f MiB/s\n", n, k, (double(n)*k)/(1024*1024)/secs.count());
}
EDIT: Also note that a big write to a pipe (bigger than PIPE_BUF) may require multiple syscalls on the read side.EDIT 2: Also, it appears that the kernel is smart enough to not copy anything when it's clear that there is no need. When I don't go through cat, I get rates that are well above memory bandwidth, implying that it's not doing any actual work:
$ ./a.out 1000000 1000 >/dev/null
buffer size: 1000000, num syscalls: 1000, perf: 1827368.373827 MiB/sI may be wrong, though. Check with lsof or similar.
Python3: 3 MiB/s
Node: 350 KiB/s
Lua: 12 MiB/s
lua -e 'while true do io.write("1") end' | pv > /dev/null
Haskell: 5 MiB/s loop = do
putStr "1"
loop
main = loop
Awk: 4.2 MiB/s yes | awk '{printf("1")}' | pv > /dev/null while true do
io.write "1"
end
PUC-Rio 5.1: 25 MiB/sPUC-Rio 5.4: 25 MiB/s
LuaJIT 2.1.0-beta3: 550 MiB/s <--- WOW
They all go slightly faster if you localize the reference to `io.write`
local write = io.write
while true do
write "1"
endNo noticeable difference for LuaJIT, which makes sense, since JIT should figure it out without help.
5.1 and 5.4 show about ~8% improvement.
main = putStr (repeat '1')
[Edit: as pointed out below, this is no longer the case!]Strings are printed one character at a time in Haskell. This choice is justified by unpredictability of the interaction between laziness and buffering; I am uncertain it's the correct choice, but the proper response is to use Text where performance is relevant.
poll([{fd=1, events=POLLOUT}], 1, 0) = 1 ([{fd=1, revents=POLLOUT}])
write(1, "11111111111111111111111111111111"..., 8192) = 8192
poll([{fd=1, events=POLLOUT}], 1, 0) = 1 ([{fd=1, revents=POLLOUT}])
write(1, "11111111111111111111111111111111"..., 8192) = 8192
With the recursive code, it buffered the output in the same way but bugged the kernel a whole lot more in-between writes. Not exactly sure what is going on: poll([{fd=1, events=POLLOUT}], 1, 0) = 1 ([{fd=1, revents=POLLOUT}])
write(1, "11111111111111111111111111111111"..., 8192) = 8192
rt_sigprocmask(SIG_BLOCK, [INT], [], 8) = 0
clock_gettime(CLOCK_PROCESS_CPUTIME_ID, {tv_sec=0, tv_nsec=920390843}) = 0
rt_sigprocmask(SIG_SETMASK, [], NULL, 8) = 0
rt_sigprocmask(SIG_BLOCK, [INT], [], 8) = 0
clock_gettime(CLOCK_PROCESS_CPUTIME_ID, {tv_sec=0, tv_nsec=920666397}) = 0
...
rt_sigprocmask(SIG_SETMASK, [], NULL, 8) = 0
poll([{fd=1, events=POLLOUT}], 1, 0) = 1 ([{fd=1, revents=POLLOUT}])
write(1, "11111111111111111111111111111111"..., 8192) = 8192I'm also not sure what's going on in the second case. IIRC, at some point historically, a sufficiently tight loop could cause trouble with handling SIGINT, so it might be related to some overagressive workaround for that?
- Node (no buffering): 1.2 MiB/s
- Go (no buffering): 2.4 MiB/s
- Python (8 KiB buffer): 2.7 MiB/s
- Go (8 KiB buffer): 218 MiB/s
Go program:
f := bufio.NewWriterSize(os.Stdout, 8192)
for {
f.WriteRune('1')
}More than a few old backup or transfer scripts had extra dd or similar tools in the pipeline to create larger and semi-asynchronous buffers, or to re-size blocks on output to something handled better by the receiver, which was a big deal on high speed tape drives back in the day. I suspect most modern hardware devices have large enough static RAM and fast processors to make that mostly irrelevant.
python actually buffers its writes with print only flushing to stdout occasionally, you may want to try:
python3 -c 'while (1): print (1, end="", flush=True)' | pv > /dev/null
which I find goes much slower (550Kib/s) yes | pv > /dev/null
There's an article floating around [1] about how yes(1) is extremely optimized considering its original purpose. In care you're wondering, yes(1) is meant for commands that (repeatedly) ask whether to proceed, expecting a y/n input or something like that. Instead of repeatedly typing "y", you just run "yes | the_command".Not sure about how yes(1) compares to the techniques presented in the linked post. Perhaps there's still room for improvement.
[1] Previous HN discussion: https://news.ycombinator.com/item?id=14542938
pv < /dev/zero > /dev/nullyes lets you specify which character to output. 'yes n' for example to output n.
yes 123abc
will print 123abc123abc123abc123abc123abc
and so on. 123abc
123abc
123abc
...Honest question: what are the practical use cases of this?
Repeatedly typing the 'y' character into a Linux pipe is surely not that common, especially at that bit rate. Also seems like the bottleneck would always be the consuming program...
At that rate no but I definitely use it once in a while. For example if a copy quite a few files and then get repeatedly asked if I want to overwrite the destination (when it's already present). Sure, I could get my commmand back and use the proper flag to "cp" or whatever to overwrite, but it's usually much quicker to just get back the previous line, go at the beginning (C-a), then type "yes | " and be done with it.
Note that you can pass a parameter to "yes" and then it repeats what you passed instead of 'y'.
It's probably still common in installation scripts, like in Dockerfiles. `apt-get install` has the `-y` option, but it would be useful for all other programs that don't.
It also allows you to script otherwise interactive command line operations with the correct answer. Many come like tools now days provide specific options to override queries. But there are still a couple hold outs which might not.
It's not made to be fast; it's just fast by nature, because there's no other computation it needs to do than to just output the string.
However, my point was just that its performance was never a main point of it. Even without optimizations, it's still very fast, and I don't think whoever created it first was concerned with it having to be super fast, as long as it was faster than the prompts of whatever was downstream in the pipe.
Using OP's code for following
php 1.8mb/sec
python 3.8 Mb/sec
node 1.0 Mb/sec
Java print 1.3 Mb/sec echo 'class Code {public static void main(String[] args) {while (true){System.out.print("1");}}}' >Code.java; javac Code.java ; java Code | pv>/dev/null
Java with buffering 57.4 Mb/sec echo 'import java.io.*;class Code2 {public static void main(String[] args) throws IOException {BufferedWriter log = new BufferedWriter(new OutputStreamWriter(System.out));while(true){log.write("1");}}}' > Code2.java ; javac Code2.java ; java Code2 | pv >/dev/nullThat manages about 7 GiB/s reusing the same buffer, or about 300 MiB/s with clearing and refilling the buffer every time
(the magic is in using java’s APIs for writing to files/sockets, which are designed for high performance, instead of using the APIs which are designed for writing to stdout)
python3 -c 'import sys
while (1): sys.stdout.write("1")'| pv>/dev/null python3 -u -c 'import sys
while (1): sys.stdout.write("1")'| pv>/dev/null
427KiB/s python3 -c 'import sys
while (1): sys.stdout.write("1")'| pv>/dev/null
6.08MiB/sUsing python 3.9.7 on macOS Monterey.
node -e '
const stdoutWrite = util.promisify(process.stdout.write).bind(process.stdout);
(async () => {
while (true) {
await stdoutWrite("1");
}
})();
' | pv > /dev/null LuaJIT 2.1.0-beta3
Using print is about 17 MiB/s luajit -e "while true do print('x') end" | pv > /dev/null
Using io.write is about 111 MiB/s luajit -e "while true do io.write('x') end" | pv > /dev/nullBash:
while :; do printf "1"; done | ./pv > /dev/null
[ 156KiB/s]
Python3 3.7.2: python3 -c 'while (1): print (1, end="")' | ./pv > /dev/null
[1,02MiB/s]
Perl 5.22.2: perl -e 'while (true) {print 1}' | ./pv > /dev/null
[3,03MiB/s]
Node.js v12.22.1: node -e 'while (1) process.stdout.write("1");' | ./pv > /dev/null
[ 482KiB/s]Perl 5.18.0 gives me 3.5 MiB per second. Perl 5.28.3, 5.30.3, and 5.34.0 gives 4 MiB per second.
perl5.34.0 -e 'while (){ print 1 }' | pv > /dev/null
For Python 3.10.4, I get about 2.8 MiB/s as you have it written, but around 5 MiB/s (same for 3.9 but only 4 MiB/s for 3.8) with this. I also get 4.8 MiB/s with 2.7: python3 -c 'while (1): print (1)' | pv > /dev/null
If I make Perl behave like yes and print a character and a newline, it has a jump of its own. The following gives me 37.3 MiB per second. perl5.34.0 -e 'while (){ print "1\n" }' | pv > /dev/null
Interestingly, using Perl's say function (which is like a Println) slows it down significantly. This version is only 7.3 MiB/s. perl5.34.0 -E 'while (1) {say 1}' | pv > /dev/null
Go 1.18 has 940 KiB/s with fmt.Print and 1.5 MiB/s with fmt.Println for some comparison. package main
import "fmt"
func main() {
for ;; {
fmt.Println("1")
}
}
These are all macports builds.Edit: I've found this consistently building multiple data processing applications over multiple years and multiple companies
To do a proper synchronous write, you'd do something like:
node -e 'const { writeSync } = require("fs"); while (1) writeSync(1, "1");' | pv > /dev/null
That gets me ~1.1MB/s with node v18.1.0 and kernel 5.4.0.Binder comes from Palm actually (OpenBinder)
What's that?
A lot of programs store sockets in /run which is typically implemented by `tmpfs`.
Edit: it’s actually “yes” that I’ve used before for generating load. I remember reading somewhere “yes” was optimized differently than the original Unix command as part of the unix certification lawsuit(s).
Long night.
"tr '\0' 1 </dev/zero | pv >/dev/null" reports 1.38GiB/s.
"yes | pv >/dev/null" reports 7.26GiB/s.
So "/dev/urandom" may not be the best source when testing performance.
UPDATE: If you mean that you want to test how fast pipes are when there is other load in the system, then I'd suggest just running a lot of stuff in the background. But I wouldn't put the process dedicated for doing something else into the pipeline you're measuring. As a matter of fact, the numbers I gave were taken with plenty of heavy processes running in the background, such as Firefox, Thunderbird, a VM with another instance of Firefox, OpenVPN, etc. etc. :)
:)
I just had 25Gb/s internet installed (https://www.init7.net/en/internet/fiber7/), and at those speeds Chrome and Firefox (which is Chrome-based) pretty much die when using speedtest.net at around 10-12Gbps.
The symptoms are that the whole tab freezes, and the shown speed drops from those 10-12Gbps to <1Gbps and the page starts updating itself only every second or so.
IIRC Chrome-based browsers use some form of IPC with a separate networking process, which actually handles networking, I wonder if this might be the case that the local speed limit for socketpair/pipe under Linux was reached and that's why I'm seeing this.
> In March 1998, Netscape released most of the code base for its popular Netscape Communicator suite under an open source license. The name of the application developed from this would be Mozilla, coordinated by the newly created Mozilla Organization https://en.wikipedia.org/wiki/Mozilla_Application_Suite#Hist...
Netscape Communicator (or Netscape 4) was released in 1997, so If we are tracing lineage, I'd say Firefox has a 2 year head start.
Firefox only uses Webkit on iOS, due to Apple requirements. It uses Gecko everywhere else. And I don't think it's ever been Safari-based anywhere.
> Firefox is only Chrome-based on iOS.
So I'm talking only about iOS. When I said it's Safari-based, I meant Webkit based, but I thought Firefox/Chrome actually pull parts of Safari on iOS. Quick research says that's wrong and they just use Webkit. Not an iOS dev, so someone can point out better sources for the 100% correct terminology.
"The gigabit has the unit symbol Gbit or Gb."
25GB/s would translate to 200Gbit/s and also 200Gb/s.
AFAIK, Firefox is not Chrome-based anywhere.
On iOS it uses whatever iOS provides for webview - as does Chrome on iOS.
Firefox and Safari is now the only supported mainstream browsers that has their own rendering engines. Firefox is the only that has their own rendering engine and is cross platform. It is also open source.
Not technically "Chrome-based", but Firefox draws graphics using Chrome's Skia graphics engine.
Firefox is not completely independent from Chrome.
In any case, i thought chrome used libnss which is a mozilla library, so you could say the reverse as well.
Interestingly safaris rendering engine is open source and cross platform, but the browser is not. Lots of linux-focused browsers (konquerer, gnome web, surf) and most embedded browsers (nintendo ds & switch, playstation) use webkit. Also some user interfaces (like WebOS, which is running all of LG's TVs and smart refrigerators) use webkit as their renderer.
https://en.wikipedia.org/wiki/File:Timeline_of_web_browsers....
https://en.wikipedia.org/wiki/Timeline_of_web_browsers has tables as well.
With it in the way I get ~15-20Gb/s
$ iperf3 -l 1M --window 64M -P10 -c speedtest.init7.net
..
[SUM] 0.00-1.00 sec 1.87 GBytes 16.0 Gbits/sec 181406
$ iperf3 -R -l 1M --window 64M -P10 -c speedtest.init7.net
..
[SUM] 0.00-1.00 sec 2.29 GBytes 19.6 Gbits/sec(Which is similar to how K8S abuses ip-tables and makes it useless for other ends, and makes you install a dedicated firewall in front of your ingress path, but let's not digress).
On the other hand, Firefox is neither chromium based, nor is a cousin of it. It's a completely different codebase, inherited from Netscape days and evolved up to this point.
As another test point, Firefox doesn't even blink at a symmetric gigabit connection going at full speed (my network is capped by my NIC, the pipe is way fatter).
FWIW Firefox under Linux (Firefox Browser 100.0.2 (64-bit)) behaves pretty much the same as Chrome. The speed raises quickly to 5-8Gb/s, then the UI starts choking, and the shown speed drops to 500Mb/s. It could be that there's some scheduling limit or other bottleneck hit in the OS itself, assuming these are different codebases (are they?).
However, you can test the limits of Linux by installing CLI version of Speedtest and hitting a nearby server.
The bottleneck maybe in the browser itself, or in your graphics stack, too.
Linux can do pretty amazing things in the network department, otherwise 100Gbps Infiniband cards wouldn't be possible at Linux servers, yet we have them on our systems.
And yes, Chrome and Firefox are way different browsers. I can confidently say this, because I'm using Firefox since it's called Netscape 6.0 (and Mozilla in Knoppix).
Mozilla suite/seamonkey isn't usually considered the same as firefox, although obviously related.
Firefox 0.8 did not have netscape branding http://theseblog.free.fr/firefox-0.8.jpg
I know. Firefox was not even an idea when Netscape 6 was released. However, inverse is true. Firefox is based on Netscape. It's just branched off actually. It started as a pared down version of SeaMonkey apparently.
The thing I was remembering from Knoppix 3.x days was "Mozilla Navigator" of SeaMonkey/Mozilla Suite, which is even older than Firefox, and discontinued 3 years later. I just booted the CD to look at it.
At the end of the day, Firefox is just Netscape Navigator, evolved.
Also, Infiniband can directly RDMA to and from MPI processes for making "remote memory local", allowing very low latencies and high performance in HPC environments.
I also like this post from Cloudflare [1]. I've read it completely, but the specifics are lost on me since I'm not directly concerned with the network part of our system.
[0]: https://medium.com/@penberg/on-kernel-bypass-networking-and-...
[1]: https://blog.cloudflare.com/how-to-receive-a-million-packets...
Totally tangential - it looks like io_uring is evolving beyond just io and into an alternate syscall interface, which is pretty neat imho.
Unfortunately the security industry has proven the why threads are a bad ideas for applications when security is a top concern.
Same applies to dynamically loaded code as plugins, where the host application takes the blame for all instabilty and exploits they introduce.
If you need security, you need isolation. If you want hardware-level isolation, you need processes. That's normal.
My disagreement with Google's applications are how they're behaving like they're the only running processes on the system itself. I'm pretty aware that some of the most performant or secure things doesn't have the prettiest implementation on paper.
I believe the default behavior is "Coalesce tabs into the same content process if they're from the same trust domain".
Then you can make it more aggressive like "Don't coalesce tabs ever" or less aggressive like "Just have one content process". I think.
I'm not sure how Firefox decides when to spawn new processes. I know they have one GPU process and then multiple untrusted "content processes" that can touch untrusted data but can't touch the GPU.
I don't mind it. It's a trade-off between security and overhead. The IPC is pretty efficient and the page cache in both Windows and Linux _should_ mean that all the code pages are shared between all content processes.
Static pages actually feel light to me. I think crappy webapps make the web slow, not browser security.
(inb4 I'm replying to someone who works on the Firefox IPC team or something lol)
The danger and joy of commenting on HN!
Even if I was working on Firefox/Chrome/whatever, I'd not be mad at someone who doesn't know something very well. Why should I? We're just conversing here.
Also, I've been very wrong here at times, and this improved my conversation / discussion skills a great deal.
So, don't worry, and comment away.
Optics: To ISP: Flexoptics (https://www.flexoptix.net/de/p-b1625g-10-ad.html?co10426=972...), Router-PC: https://mikrotik.com/product/S-3553LC20D
Router: Mikrotik CCR-2004 - https://mikrotik.com/product/ccr2004_1g_12s_2xs - warning: it's good to up to ~20Gb/s one way. It can handle ~25Gb/s down, but only ~18Gb/s up, and with IPv6 the max seems to be ~10Gb/s any direction.
If Mikrotik is something you're comfortable using you can also take a look at https://mikrotik.com/product/ccr2216_1g_12xs_2xq - it's more expensive (~2500EUR), but should handle 25Gb/s easily.
Turned out windows 7 or the NICs needed a lot of tuning to work well. There was alot of freezing and other fail.
But alas, they do not, lol. They just use no-cache on the query which will not prevent the browser from storing the data.
Oh yes, Linux pipes were invented by Douglas McIlroy while working for Bell Labs on Research UNIX and first described in the man pages of Version 3 Unix, Feb. 1974, just a couple months after Linus Torvald's 4th birthday.
Where and how and when will the unjust and blatent plagiarism of Linux cease? The software was made free by BSD, so feel free to use it, roll it all into GNU/Linux, have at it, but please stop incorrectly describing these things as Linux things. Because the only software that I am certain actually belongs to Linux is systemd. So let's start calling that "Linux systemd," and stop calling anything else Linux anything.
Seems to me like this post is pretty specifically about Unix pipes on Linux (i.e. Linux pipes), as opposed to Unix pipes in general.
The article also talks about “Linux paging”, again clearly referring to the implementation and usage of virtual memory on Linux, rather than…whatever ancient architecture first invented the page table.
So I'm not sure it is not analogous to, "today, we're going to talk about Linux electricity, specifically the way electricity is utilized in Linux."