We have had very similar issues with Java - excessive heap allocations can bring the system to a crawl and there is often no easy way out.
We have had very similar issues with Java - excessive heap allocations can bring the system to a crawl and there is often no easy way out.
$ pv lol.tsv | node lol.js
1.71GiB 0:00:07 [ 225MiB/s] [================================>] 100%
Finished reading the file.
Do we live in a world where engineers using a mature language like NodeJS with a pretty simple line-by-line file reading process and a string split think they hit a hard wall at 700kb/s, and the problem is not their bad code?1. https://gist.github.com/orf/92bc26381e3dc00c8e96af262768591b
The point of the article is that you can write really boring normal code in NodeJS and it's really slow, or you can write really boring normal code in Rust and it's a lot faster, and it's not that hard to use the latter to replace hot paths in the former.
I know basically nothing about Node though, so I could be completely wrong here.
If they are seeing any GC activity once the loop gets hot, something else is going very wrong.
1. It should never be that stupidly slow
2. That should be obvious to anyone with a napkin and some math
3. Excessive allocations are highly unlikely to be the root cause of the slowness
4. Instead of debugging, they seemingly spent a lot of time putting it on Kubernetes, scaling it out to 25 instances and then finally rewriting a portion of it in Rust.
The whole thing smells of confusion and multiple issues being conflated together, with the rewrite in rust fixing these by virtue of deletion.
And if the root is was the GC, then some `--trace-gc` output in the post would be nice at the absolute minimum. Showing a for loop with a string split doesn't cut it.
And I can definitely say: IO speed is pretty good. To be honest, I've never come across a language with bad IO speed and I've seen quite a lot of languages.
I also did some benchmarks (like orf did) and speed is pretty good. Maybe Rust is a little bit better because it's nearer to syscalls than Node.js is but we're not talking about 760kb/s vs. 250mb/s.
200GB of logs and 25 instances means 8GB per instance. Not sure what they're doing, but processing of 8GB taking 3h is like doing it with my 80286 processor ;-) Even with inserting stuff in MongoDB this seems way to slow. And they only replaced the file IO part with Rust...
I guess the records array is getting pretty huge which might result in swapping to disk. But then: don't let it happen, just flush it to the database or whatever. I believe the problem could have been solved with Node.js without issues.
With the limited information provided in the article, there are other points that I would check before going as far as writing a module in another language. For example
> records.push({
They are collecting all the parsed rows in a single array. This is going to consume memory proportional to the size of the file, potentially in the same order of magnitude (depends on how much information is being collected and how much discarded).
In my experience, a common mistake in handling files too large for memory is to read the data faster than you process it, consequently running out of memory.
For example, maybe they have a loop reading lines and doing a bunch of async DB writes: your process will read at disk speed and write at database speed (relatively slower), and you'll run out of memory due to async tasks/context queuing up in memory.
Perhaps this is why it worked to split the file into 8GB units X 25 servers: it didn't blow their memory budget per server.
With my first try with bigger data it blew up with an out of heap memory error. The article mentions them having memory problems as well. They probably operated at the upper levels of available memory and even got into swapping and then 700kb/s is not surprising.
So yes, this absolutely looks like a memory allocation issue.
Try modifying it to clear the array (a DB flush, if you will) every 100k lines.
Now we don't have any Rust code to compare so we can't say how it's different. But everything points towards a memory allocation issue that is just not their in simple Rust code.
* "unbounded array" - are you referring to `records`?
* If so, is the memory of `records` too large? Or is it that `records[i].pathname` has a reference to `fields[7]`, AND v8 will retain all copies of `fields` until the loop is done?
Do you think this code would be more GC-friendly?
```
let count = 0;
const MAX = 1000;
const flushedRecords = [];
for await (const line of readlineStream) {
const fields = line.split('\t');
const httpStatus = Number(fields[5]);
if (httpStatus < 200 || httpStatus > 299) continue;
records.push({
pathname: fields[7],
referrer: fields[8],
// ...
});
if(count == MAX || isLastLine(line)){
flushedRecords.push([...records]);
records = [];
count = 0;
} else {
count++;
}
}```
Or were you thinking something totally different?
1. It parses 200gb at an average rate of 700kb/s
2. The job splits strings into a format that is loaded into mongodb
3. There where OOM errors that went away with a smaller workload
None of this points to an inherent issue with NodeJS and the way it references memory.
The speed issue alone points to a different problem root cause, and the post is completely lacking in concrete analysis.
If v8 is being dumb and retains references to huge strings, then understanding this is the case and forcing a copy (and thus enabling the gc to reclaim data) is quicker and easier than any of their attempts at fixing the issue.
Splitting a small TSV file and loading it into mongo can be done with even the most memory inefficient language without issue.
I do agree here, but I think it's telling that the rust solution is, without analysis, just better.
Is this the best developer or best case study? Perhaps not, but it is interesting in that little analysis was needed.
In the case where you have some constraint placed on you (running in the browser, company policy forbids rust e.g.), the analysis is worthwhile. Otherwise, assuming the fix didn't take long (relative to the requisite analysis and dev training), the decision to switch to a less shoot-yourself-in-the-foot language seems reasonable to me.
BUT, what I am interested in - if I'm using javascript to parse a large file line by line, plucking off specific fields like the example does, what can I try to see if it's more GC-friendly? Obviously the answer is always some flavor of "profile it" and "it depends". If the blog author could have gotten a big boost in performance by adding 10 lines of js code to their loop, that is something all readers of the article could benefit from.
I will play around with it and see what I can learn about the V8 GC. For reference: here is a talk about profiling the ruby gc, which helped me fix a memory leak at work (https://www.youtube.com/watch?v=kZcqyuPeDao).
* How long does it take to ramp up on the new code base (author notwithstanding as a user of Rust outside of the company)?
* If the author of the code is sick and there's a bug, or logic needs to be added, how hard would it be for someone else to do this work?
* What is the build story for the Rust code like? How does it integrate with existing CI?
* How do you store the Rust library? Some artifact store that Node can interact with?
Without more details, unless the company was already writing Rust, this doesn't seem like a reasonable choice at all. Doing some profiling and experimenting with allocation, along with comments as to why you're doing something that's not naive (e.g. appending to a large array) seems much cheaper than doing all this Rust work.
I've been in situations at $WORK where we were locked into using Python for certain code. Rewriting portions of it in Rust (if it needed running in the runtime) or Java (if it was across a service boundary) would have been enjoyable. But answering all those questions made us stick to Python. And surely, in most of those instances, we never really came back to those hot loops again so it never needed subsequent work. There were a few cases where we did find ourselves constantly trying to eke out more performance and in most of those cases a rewrite did end up going through, but those are rare and part of the craft of software engineering is to identify these areas and get ahead of rewrites early enough to not waste cycles.
IMO that's not an inappropriate response here. You can make this code in Node, but it's super-awkward trying to reduce string allocations in JavaScript. Much easier to rewrite in Rust where such things will be fast even with the naive code. For code such as this you can likely do a 1:1 port of the Node code and it will be 10x faster. Which is much better than getting yourself in twists trying to make Node fast at tasks it's not good at.
Does this really, really not scream “there is another issue” to you?
Do you really think, even with string allocations in a hot loop, NodeJS with its very optimised runtime would struggle to hit 1mb/s?
This channel is a real gem, show what talent really is: https://www.youtube.com/watch?v=a1L7k35EHIc
Most of your application probably doesn't need it and the tight loop might not be big enough to be worth using JNI to write in a non-GC language (as well as maintaining another compiler stack for)
Ideally the runtime should allow you to manually manage memory when you need to, like some keyword that lets you manually keep track of an object lifetime. But as far as I know most GC languages don't have that
Anyway, I'm not sure if this is in Java already or is upcoming, but Java will finally add value objects, that is, stack variables instead of just making everything a heap object. For tight loops that will be much more efficient, since it can just reuse the same memory, stay within the cpu registers, etc. (I'm talking out of my ass btw from a superficial knowledge of Go and Java and half-remembered theory from college).
Does node not similarly allocate on the stack in hot loops?